Skip to content

feat(md): Zig markdown parser with Bun.markdown API - #26440

Merged
dylan-conway merged 56 commits into
mainfrom
jarred/mdx
Jan 29, 2026
Merged

dylan-conway merged 56 commits into
mainfrom
jarred/mdx

Conversation

@Jarred-Sumner

@Jarred-Sumner Jarred-Sumner commented Jan 25, 2026 •

Copy link
Copy Markdown
Collaborator

Summary

  • Port md4c (CommonMark-compliant markdown parser) from C to Zig under src/md/
  • Three output modes:
    • Bun.markdown.html(input, options?) — render to HTML string
    • Bun.markdown.render(input, callbacks?) — render with custom callbacks for each element
    • Bun.markdown.react(input, options?) — render to a React Fragment element, directly usable as a component return value
  • React element creation uses a cached JSC Structure with putDirectOffset for fast allocation
  • Component overrides in react(): pass tag names as options keys to replace default HTML elements with custom components
  • GFM extensions: tables, strikethrough, task lists, permissive autolinks, disallowed raw HTML tag filter
  • Wire up .md as a bundler loader (via explicit { type: "md" })

JavaScript API

Bun.markdown.html(input, options?)

Renders markdown to an HTML string:

const html = Bun.markdown.html("# Hello **world**");
// "<h1>Hello <strong>world</strong></h1>\n"

Bun.markdown.html("## Hello", { headingIds: true });
// '<h2 id="hello">Hello</h2>\n'

Bun.markdown.render(input, callbacks?)

Renders markdown with custom JavaScript callbacks for each element. Each callback receives children as a string and optional metadata, and returns a string:

// Custom HTML with classes
const html = Bun.markdown.render("# Title\n\nHello **world**", {
  heading: (children, { level }) => `<h${level} class="title">${children}</h${level}>`,
  paragraph: (children) => `<p>${children}</p>`,
  strong: (children) => `<b>${children}</b>`,
});

// ANSI terminal output
const ansi = Bun.markdown.render("# Hello\n\n**bold**", {
  heading: (children) => `\x1b[1;4m${children}\x1b[0m\n`,
  paragraph: (children) => children + "\n",
  strong: (children) => `\x1b[1m${children}\x1b[22m`,
});

// Strip all formatting
const text = Bun.markdown.render("# Hello **world**", {
  heading: (children) => children,
  paragraph: (children) => children,
  strong: (children) => children,
});
// "Hello world"

// Return null to omit elements
const result = Bun.markdown.render("# Title\n\n![logo](img.png)\n\nHello", {
  image: () => null,
  heading: (children) => children,
  paragraph: (children) => children + "\n",
});
// "Title\nHello\n"

Parser options can be included alongside callbacks:

Bun.markdown.render("Visit www.example.com", {
  link: (children, { href }) => `[${children}](${href})`,
  paragraph: (children) => children,
  permissiveAutolinks: true,
});

Bun.markdown.react(input, options?)

Returns a React Fragment element — use it directly as a component return value:

// Use as a component
function Markdown({ text }: { text: string }) {
  return Bun.markdown.react(text);
}

// With custom components
function Heading({ children }: { children: React.ReactNode }) {
  return <h1 className="title">{children}</h1>;
}
const element = Bun.markdown.react("# Hello", { h1: Heading });

// Server-side rendering
import { renderToString } from "react-dom/server";
const html = renderToString(Bun.markdown.react("# Hello **world**"));
// "<h1>Hello <strong>world</strong></h1>"

React 18 and older

By default, react() uses Symbol.for('react.transitional.element') as the $$typeof symbol, which is what React 19 expects. For React 18 and older, pass reactVersion: 18:

const el = Bun.markdown.react("# Hello", { reactVersion: 18 });

Component Overrides

Tag names can be overridden in react():

Bun.markdown.react(input, {
  h1: MyHeading,      // block elements
  p: CustomParagraph,
  a: CustomLink,      // inline elements
  img: CustomImage,
  pre: CodeBlock,
  // ... h1-h6, p, blockquote, ul, ol, li, pre, hr, html,
  //     table, thead, tbody, tr, th, td,
  //     em, strong, a, img, code, del, math, u, br
});

Boolean values are ignored (not treated as overrides), so parser options like { strikethrough: true } don't conflict with component overrides.

Options

Bun.markdown.html(input, {
  tables: true,              // GFM tables (default: true)
  strikethrough: true,       // ~~deleted~~ (default: true)
  tasklists: true,           // - [x] items (default: true)
  headingIds: true,          // Generate id attributes on headings
  autolinkHeadings: true,    // Wrap heading content in <a> tags
  tagFilter: false,          // GFM disallowed HTML tags
  wikiLinks: false,          // [[wiki]] links
  latexMath: false,          // $inline$ and $$display$$
  underline: false,          // __underline__ (instead of <strong>)
  // ... and more
});

Architecture

Parser (src/md/)

The parser is split into focused modules using Zig's delegation pattern:

Module Purpose
parser.zig Core Parser struct, state, and re-exported method delegation
blocks.zig Block-level parsing: document processing, line analysis, block start/end
containers.zig Container management: blockquotes, lists, list items
inlines.zig Inline parsing: emphasis, code spans, HTML tags, entities
links.zig Link/image resolution, reference links, autolink rendering
autolinks.zig Permissive autolink detection (www, url, email)
line_analysis.zig Line classification: headings, fences, HTML blocks, tables
ref_defs.zig Reference definition parsing and lookup
render_blocks.zig Block rendering dispatch (code, HTML, table blocks)
html_renderer.zig HTML renderer implementing Renderer VTable
types.zig Shared types: Renderer VTable, BlockType, SpanType, TextType, etc.

Renderer Abstraction

Parsing is decoupled from output via a Renderer VTable interface:

pub const Renderer = struct {
    ptr: *anyopaque,
    vtable: *const VTable,

    pub const VTable = struct {
        enterBlock: *const fn (...) void,
        leaveBlock: *const fn (...) void,
        enterSpan:  *const fn (...) void,
        leaveSpan:  *const fn (...) void,
        text:       *const fn (...) void,
    };
};

Four renderers are implemented:

  • HtmlRenderer (src/md/html_renderer.zig) — produces HTML string output
  • JsCallbackRenderer (src/bun.js/api/MarkdownObject.zig) — calls JS callbacks for each element, accumulates string output
  • ParseRenderer (src/bun.js/api/MarkdownObject.zig) — builds React element AST with MarkedArgumentBuffer for GC safety
  • JSReactElement (src/bun.js/bindings/JSReactElement.cpp) — C++ fast path for React element creation using cached JSC Structure + putDirectOffset

Test plan

  • 792 spec tests pass (CommonMark, GFM tables, strikethrough, tasklists, permissive autolinks, GFM tag filter, wiki links, coverage, regressions)
  • 114 API tests pass (html(), render(), react(), renderToString integration, component overrides)
  • 58 GFM compatibility tests pass
bun bd test test/js/bun/md/md-spec.test.ts       # 792 pass
bun bd test test/js/bun/md/md-render-api.test.ts  # 114 pass
bun bd test test/js/bun/md/gfm-compat.test.ts     # 58 pass

🤖 Generated with Claude Code

Port md4c (CommonMark-compliant markdown parser) from C to Zig under
src/md/. Wire up .md as a loader that outputs HTML, and expose
Bun.Markdown.renderToHTML() for programmatic use with per-call options.

Parser passes all 768 spec tests: CommonMark, GFM tables, strikethrough,
tasklists, permissive autolinks, wiki links, coverage, and regressions.

Co-Authored-By: Claude <noreply@anthropic.com>
@robobun

robobun commented Jan 25, 2026 •

Copy link
Copy Markdown
Collaborator
Updated 9:09 PM PT - Jan 28th, 2026

❌ @dylan-conway, your commit 7bd8073 has 3 failures in Build #36054 (All Failures):


🧪   To try this PR locally:

bunx bun-pr 26440

That installs a local version of the PR into your bun-26440 executable, so you can run:

bun-26440 --bun

autofix-ci Bot and others added 4 commits January 25, 2026 09:36
…nder() API

- Remove .md/.markdown default loader association (explicit {type: "md"} still works)
- Split monolithic parser.zig (~4500 lines) into 8 focused modules using
  Zig's delegation pattern: blocks, containers, inlines, links, autolinks,
  line_analysis, ref_defs, render_blocks
- Introduce Renderer VTable abstraction to decouple parsing from output,
  enabling pluggable renderers (HTML, JS callbacks, etc.)
- Add Bun.Markdown.render(input, callbacks) JS API with 21 callback types
  (heading, paragraph, strong, emphasis, link, image, code, codespan,
  blockquote, list, listItem, hr, table, thead, tbody, tr, th, td, html,
  strikethrough, text) using content-stack pattern for nested rendering
- Add GFM-specific spec tests (spec-gfm.txt) covering autolinks, code
  spans, entity references, links, emphasis interactions

Co-Authored-By: Claude <noreply@anthropic.com>
Add tag_filter option that replaces the leading `<` with `&lt;` for 9
disallowed HTML tags (title, textarea, style, xmp, iframe, noembed,
noframes, script, plaintext) per GFM spec section 6.11. The filter
applies to both HTML blocks and inline HTML spans, with case-insensitive
matching. Enabled in the github options preset.

Co-Authored-By: Claude <noreply@anthropic.com>
@remorses

remorses commented Jan 25, 2026 •

Copy link
Copy Markdown
Contributor

It would be cool if there was a method Markdown.parse(string) that returned MDAST compatible AST.

This way you could reuse it with remark plugins. Very commonly used to transform markdown before rendering.

Here is MDAST spec: https://github.com/syntax-tree/mdast

Here is an example of how MDAST looks like

@Jarred-Sumner
Jarred-Sumner marked this pull request as ready for review January 25, 2026 22:19
@Jarred-Sumner Jarred-Sumner changed the title feat(md): add Zig markdown parser with .md loader and Bun.Markdown API feat(md): Zig markdown parser with Bun.Markdown API Jan 25, 2026
…erer tests

- Rename Bun.Markdown → Bun.markdown (lowercase)
- Rename renderToHTML → html
- Add custom HTML renderer test (reimplements HTML from callbacks)
- Add ANSI terminal renderer test (headings, bold, italic, links,
  code blocks with borders, blockquotes with bar prefix, bullet lists)

Co-Authored-By: Claude <noreply@anthropic.com>
@coderabbitai

coderabbitai Bot commented Jan 25, 2026 •

Copy link
Copy Markdown
Contributor

Caution

Review failed

The pull request is closed.

Walkthrough

Adds a Markdown parsing and rendering subsystem (Bun.markdown) with GFM features, JSON5 and .md loader support across bundler/transpiler/runtime, React element bindings for render-to-JSX, TypeScript types, tests, and docs; numerous new md/* Zig modules implement parsing, rendering, and Unicode helpers.

Changes

Cohort / File(s) Summary
Markdown Core Modules
src/md/root.zig, src/md/parser.zig, src/md/types.zig, src/md/helpers.zig, src/md/unicode.zig, src/md/blocks.zig, src/md/containers.zig, src/md/line_analysis.zig, src/md/inlines.zig, src/md/links.zig, src/md/ref_defs.zig, src/md/render_blocks.zig, src/md/html_renderer.zig, src/md/autolinks.zig, src/md/entity.zig
New, large Markdown parser/renderer implementation: parser, block/inline processing, reference defs, link/autolink handling, render pipelines (HTML/React/callback), Unicode case-folding and helpers, entity decoding, and public API surface.
Runtime API & Bindings
src/bun.js/api.zig, src/bun.js/api/BunObject.zig, src/bun.js/api/MarkdownObject.zig, src/bun.js/bindings/BunObject+exports.h, src/bun.js/bindings/BunObject.cpp
Expose Bun.markdown namespace; add MarkdownObject with create/render/React entry points; register lazy BunObject.markdown property and corresponding C bindings.
React Element Bindings
src/bun.js/bindings/JSReactElement.h, src/bun.js/bindings/JSReactElement.cpp, src/bun.js/bindings/ZigGlobalObject.h, src/bun.js/bindings/ZigGlobalObject.cpp
Add JSReactElement structure, factory C APIs (create/createFragment), global-object integration, and late init of React element structure.
Loader & Bundler Integration
src/api/schema.zig, src/options.zig, src/bun.js/ModuleLoader.zig, src/bun.js/bindings/ModuleLoader.cpp, src/bun.js/bindings/headers-handwritten.h, src/bundler/LinkerContext.zig, src/bundler/ParseTask.zig, src/js_printer.zig, src/transpiler.zig, src/bun.zig, cmake/Sources.json
Add .md and .json5 loader support across schema/options/ModuleLoader/transpiler/parse task/js_printer; JSON5 parsing path and .md → renderToHtml integration; update Loader enum and mappings; minor formatting change in Sources.json.
Bundler/DevServer Handling
src/bake/DevServer/DirectoryWatchStore.zig, src/bun.js/ModuleLoader.zig
Treat .json5 and .md in resolution-failure handling and loader switch paths consistent with other non-JS assets.
Bun Core Exports
src/bun.zig
Add public import alias md for ./md/root.zig.
TypeScript Definitions
packages/bun-types/bun.d.ts, packages/bun-types/jsx.d.ts
Add comprehensive Bun.markdown typings (Options, callbacks, React types) and JSX.Element conditional type; note: Markdown namespace block duplicated in bun.d.ts.
Tests — Suites & Specs
test/js/bun/md/md-spec.test.ts, test/js/bun/md/gfm-compat.test.ts, test/js/bun/md/md-edge-cases.test.ts, test/js/bun/md/md-react.test.ts, test/js/bun/md/md-render-callback.test.ts, test/js/bun/md/md-heading-ids.test.ts, test/js/bun/md/*spec-*.txt, test/js/bun/md/*
Add extensive unit/spec/regression/compatibility tests and spec files for CommonMark/GFM features, edge cases, coverage, and regressions.
Docs
docs/docs.json, docs/runtime/markdown.mdx, docs/runtime/bun-apis.mdx
Add runtime docs for Bun.markdown and JSON5 pages; update API reference to include Bun.markdown.
Bundler AST / Transpiler
src/bundler/ParseTask.zig, src/transpiler.zig, src/js_printer.zig
JSON5 parsing into AST; Markdown (.md) rendering at transpile time to HTML string modules; emit loader metadata for md.
Test Harness / Types Tests
test/integration/bun-types/bun-types.test.ts
Refactor TypeScript test harness to isolated fixtures and helper utilities; various internal test adjustments.
Misc Tests & Data
test/js/bun/md/coverage.txt, test/js/bun/md/regressions.txt
Add coverage and regression test data files.
Small edits
src/bake/DevServer/DirectoryWatchStore.zig, src/bun.js/bindings/headers-handwritten.h, src/bun.js/bindings/ModuleLoader.cpp, src/bun.js/ModuleLoader.zig
Add loader constants and map "md" to BunLoaderTypeMD; update error messaging.

Possibly related PRs

Suggested reviewers

  • alii
  • lydiahallie
  • pfgithub
🚥 Pre-merge checks | ✅ 2
✅ Passed checks (2 passed)
Check name Status Explanation
Title check ✅ Passed The title accurately describes the main change: adding a Zig markdown parser with Bun.markdown API, which aligns with the extensive changes across parser modules, bindings, and documentation.
Description check ✅ Passed The description comprehensively covers the PR scope with summary, three API modes, architecture overview, test plan, and code examples, exceeding template requirements by providing detailed documentation.

✏️ Tip: You can configure your own custom pre-merge checks in the settings.


Comment @coderabbitai help to get the list of available commands and usage tips.

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 25

🤖 Fix all issues with AI agents
In `@src/bun.js/api/MarkdownObject.zig`:
- Around line 121-160: Rename the module-private struct fields in
JsCallbackRenderer (and its nested Callbacks and StackEntry) to use a leading
'#' (e.g., globalObject -> `#globalObject`, allocator -> `#allocator`, src_text ->
`#src_text`, stack -> `#stack`, callbacks -> `#callbacks`, has_js_error ->
`#has_js_error`, and similar for all fields inside Callbacks and StackEntry) and
update all references/call sites to the new private names (search for usages of
JsCallbackRenderer, Callbacks, and StackEntry and replace field accesses
accordingly) so the structs follow the private field naming convention.
- Around line 196-275: appendToTop and the various stack.append calls currently
swallow allocation failures; change their catches to set an allocation error
flag (reuse or add a field like has_js_error or has_alloc_error on
JsCallbackRenderer) and return immediately instead of ignoring the error.
Specifically, replace the empty catches in appendToTop
(top.buffer.appendSlice(... ) catch {}) and the stack.append(...) calls in
enterBlockImpl and enterSpanImpl (and any other stack.append uses) with catches
that set self.has_js_error = true (or self.has_alloc_error = true) and return;
ensure render() (or the top-level caller) checks that flag and aborts/surfaces
an error when set so OOMs are not silently dropped.

In `@src/bundler/ParseTask.zig`:
- Around line 371-385: The markdown-to-HTML error should not be converted into
error.ParserError; change the catch on bun.md.renderToHtml in the .md arm to a
catch |err| block that logs the failure (using log.addError with the original
err context), calls bun.handleOom(err) to let OOMs be handled, and then returns
err to propagate the real error to the caller instead of returning
error.ParserError; update the code around bun.md.renderToHtml, the catch block,
and the surrounding control flow in ParseTask.zig so the original error is
returned.

In `@src/md/autolinks.zig`:
- Around line 145-251: The function findPermissiveAutolink reads content[pos]
without bounds checking; add an early guard at the top of findPermissiveAutolink
to return a not-found AutolinkResult when content is empty or pos is out of
range (e.g., if content.len == 0 or pos >= content.len) so external callers
cannot trigger an out-of-bounds access; update the initial checks in
findPermissiveAutolink (referencing pos and content) to return .{ .found =
false, .beg = 0, .end = 0 } when the guard fails.

In `@src/md/blocks.zig`:
- Around line 701-740: In endCurrentBlock, the call to
self.block_bytes.appendSlice(self.allocator, line_bytes) currently swallows
allocation failures; change it to propagate the error by checking the result of
appendSlice and returning the OutOfMemory error (or bubbling the specific error)
instead of ignoring it so callers of endCurrentBlock see allocation failures;
locate the appendSlice call in endCurrentBlock (uses self.current_block_lines
and line_bytes) and replace the catch {} with proper error handling that returns
the error from appendSlice.
- Around line 29-554: In analyzeLine, the paragraph-interrupt check uses "off >=
self.size" but then calls self.ch(cont_result.off) which accesses
cont_result.off; change the guard to check cont_result.off (i.e., use
cont_result.off >= self.size or helpers.isNewline(self.ch(cont_result.off)) as
appropriate) so the bounds check matches the actual index passed to self.ch;
update the conditional around the "Blank after list mark can't interrupt
paragraph" branch that currently references off to instead reference
cont_result.off to avoid out-of-range access when cont_result.off == self.size.

In `@src/md/containers.zig`:
- Around line 89-98: Update the misleading comment inside the
isContainerCompatible function: change the comment "Bullet lists: different
bullet chars are compatible" to accurately state that different bullet
characters are incompatible (e.g., "Bullet lists: different bullet chars are
incompatible"), so the comment matches the existing logic that returns false
when isListBullet(existing.ch) and isListBullet(new.ch); reference the
isContainerCompatible function and the isListBullet check to locate the line to
edit.

In `@src/md/helpers.zig`:
- Around line 80-175: Revert/remove the added Unicode-version comment and keep
the existing source attribution in the isUnicodePunctuationExtended function: do
not introduce a specific Unicode version note for the ranges; instead retain the
existing comment that they "match md4c's punct map" (referencing
isUnicodePunctuationExtended and its ranges array) so the code continues to
point maintainers to md4c as the authoritative source.

In `@src/md/inlines.zig`:
- Around line 576-581: The subtraction-assignment `closer_idx -%= 1` in the
while loop (used with `closer_idx` and `delims`) is a wrapping trick to undo the
loop's upcoming increment so the same closer gets re-checked; add a brief inline
comment next to that statement explaining that intent (mention `closer_idx`,
`delims[closer_idx].remaining` and `delims[closer_idx].can_close`) so future
readers understand this wrap-around subtraction is purposely undoing the loop
increment to re-process the same closer.

In `@src/md/line_analysis.zig`:
- Around line 19-21: The isHrLine function reads self.text[off] without
verifying the index is in range, mirroring the same bounds issue noted for
isSetextUnderline; add a guard to ensure off is within self.text.len (or use the
existing safe accessor) before accessing self.text[off], returning false if
out-of-bounds, and apply the same pattern used in isSetextUnderline so isHrLine
(and any similar functions) do not perform unchecked indexing.
- Around line 311-374: The function isTableUnderline mutates parser state
(self.table_alignments and self.table_col_count) despite an `is*` name that
implies a pure predicate; either rename the function to reflect side effects
(e.g., parseTableUnderline or consumeTableUnderline) or explicitly document that
it updates parser state, and update all call sites and tests to use the new name
or expect the side effects; specifically change references to isTableUnderline,
and ensure the behavior around self.table_alignments (the alignment assignments)
and self.table_col_count (final assignment) is preserved when renaming or adding
documentation.
- Around line 1-3: The function isSetextUnderline currently accesses
self.text[off] without ensuring off < self.size; add a defensive bounds check at
the start of isSetextUnderline to verify off is within range (e.g., if (off >=
self.size) return .{ .is_setext = false, .level = 0 } ) or alternatively assert
the precondition with a clear message; ensure you reference the same symbols
(isSetextUnderline, self.text, self.size, off) so the guard prevents
out-of-bounds reads before the existing character comparison.

In `@src/md/links.zig`:
- Around line 1-2: The parameter base_off on the function processLink (type
Parser) is unused and currently discarded; either remove base_off from
processLink's signature and update all call sites accordingly, or keep it but
document its intent and stop assigning it to _ (replace the discard with a
clarifying comment like "reserved for future offset handling") so readers know
it's intentionally unused; adjust any references to Parser.processLink to match
the new signature if you remove it.
- Around line 410-413: Replace the hardcoded 100-character check for wiki link
targets with a configurable limit: introduce a constant or parser field (e.g.,
WikiLinkMaxLen or Parser.wikiLinkMaxLen) and use that instead of the literal in
the validation block that checks target.len (the code containing "if (target.len
> 100)"). Ensure the new config has a sensible default of 100, propagate the
field into any functions that need it (or read the constant from the
parser/context where the link parsing function is declared), and update any
tests or callers to use the new configurable value.

In `@src/md/parser.zig`:
- Around line 52-54: The table_alignments buffer is only 64 entries but the
parser supports up to types.TABLE_MAXCOLCOUNT (128), which can cause OOB writes;
update the declaration of table_alignments to use types.TABLE_MAXCOLCOUNT (or
the TABLE_MAXCOLCOUNT constant) instead of the hardcoded 64, or alternatively
clamp uses of table_col_count/table indexing to the buffer length wherever code
references table_alignments; make the change around the table_alignments
declaration and any indexing sites that use table_col_count to ensure capacity
and bounds are consistent.
- Around line 223-235: The HtmlRenderer allocated by HtmlRenderer.init in
renderToHtml may leak because html_renderer.deinit() isn't guaranteed to run on
errors; add a defer html_renderer.deinit(); immediately after var html_renderer
= HtmlRenderer.init(allocator, input, tag_filter); (before initializing Parser
and definitely before any try calls) so the ArrayListUnmanaged inside
html_renderer is always freed on error or return, mirroring the existing defer
parser.deinit().

In `@src/md/ref_defs.zig`:
- Around line 284-300: Current code does an explicit loop over ref_defs.items to
check for an existing label then later may call lookupRefDef elsewhere, causing
redundant linear scans; replace the manual duplicate check with a single
lookupRefDef(self, norm_label) call and only proceed to allocator.dupe and
self.ref_defs.append(...) when lookupRefDef reports "not found" (or equivalent
false/err_not_found), thereby eliminating the extra iteration and centralizing
existence logic in lookupRefDef; keep the normalization value norm_label and use
it when appending .label to the new entry.
- Around line 9-48: normalizeLabel currently swallows allocation failures by
returning the original raw slice (using "catch return raw" at calls to
result.append and result.appendSlice), which leads to inconsistent behavior;
change normalizeLabel to propagate allocation errors instead: update its
signature to return an error union (e.g., ![]const u8 or []const
u8!AllocatorError), replace each "catch return raw" with "catch |err| return
err" (or propagate the allocator error), ensure you release any partially-built
ArrayListUnmanaged if needed before returning an error, and update all callers
of normalizeLabel (and any tests) to handle the new error-returning contract;
locate fixes around the normalizeLabel function and the result.append /
result.appendSlice calls as well as the UTF-8 handling using helpers.decodeUtf8,
helpers.encodeUtf8, and unicode.caseFold.
- Around line 51-59: lookupRefDef currently does linear scans over
self.ref_defs.items causing O(n²) behavior when called repeatedly (e.g., from
buildRefDefHashtable and link resolution); replace this with an O(1) hash lookup
by adding a hashtable field to Parser (or a std.HashMap instance) keyed by the
normalized label, update buildRefDefHashtable to normalize labels and populate
that map (instead of relying on repeated scans), and change lookupRefDef to
early-return via the map lookup (still validating empty/whitespace labels and
normalizing using normalizeLabel); reference functions/fields: lookupRefDef,
buildRefDefHashtable, normalizeLabel, and self.ref_defs.items to locate where to
add/replace logic.

In `@src/md/render_blocks.zig`:
- Around line 76-107: processTableRow currently skips over backtick code spans
(the '`' handling: bt_count, found_close, close_count logic) when scanning
row_text for '|' so pipes inside backticks are ignored; remove that
backtick-skip branch and let the main loop treat '`' like any other character
(keeping the existing backslash-escape handling that advances end by 2) so
unescaped '|' inside backtick spans will split cells per GFM-compat tests;
update any related comments and ensure variables
bt_count/close_count/found_close are removed or unused code cleaned up.

In `@src/transpiler.zig`:
- Around line 1358-1390: The catch on bun.md.renderToHtml currently drops the
renderer error; change the catch to capture the error (e.g., catch |err| ...)
and pass its message/details into transpiler.log.addErrorFmt so the log includes
the actual renderer error; update the transpiler.log.addErrorFmt call used in
the .md case (the block that constructs js_ast.Expr and returns ParseResult) to
include the captured err information in the formatted message while preserving
the existing context and return null behavior.

In `@test/js/bun/md/gfm-compat.test.ts`:
- Around line 19-25: The GFM renderer function renderGFM currently doesn't set
the tag_filter option, causing tagfilter-dependent tests to rely on defaults;
update the renderGFM options object to explicitly include tag_filter: true so
tag filtering is enabled (modify the renderGFM function where render(...) is
called and add the tag_filter option alongside tables, strikethrough, tasklists,
and permissive_autolinks).

In `@test/js/bun/md/md-spec.test.ts`:
- Around line 234-241: The loop over specFiles silently swallows parse errors by
using catch { continue; } which hides missing or invalid spec files; change the
catch to capture the error (e.g., catch (err)) and either rethrow a new Error
that includes the specPath and original error, or explicitly log a clear skip
reason before continuing so parseSpecFile failures are visible; update the block
around parseSpecFile(specPath) and the examples variable handling to include the
error variable and propagate or report it rather than silently continuing.

In `@test/js/bun/md/spec-permissive-autolinks.txt`:
- Around line 40-176: The spec prose contains several spelling mistakes that
need simple corrections: replace "formost" with "foremost" and all occurrences
of "alhanumerical" with "alphanumerical" (and any inconsistent uses), and fix
"Once again any"→"Once again, any" where missing commas for clarity; update
these words in the text blocks that include the phrases "First and formost",
"alhanumerical characters", and "alhanumerical" within the URL/path/query
descriptions to preserve examples and formatting.

In `@test/js/bun/md/spec-strikethrough.txt`:
- Around line 38-40: Fix the grammar in the explanatory sentence in
spec-strikethrough.txt: replace the incorrect word "then" with "they" so the
sentence reads "and similarly, they cannot close the span if they are preceded
with a whitespace." Locate the sentence in the file (the paragraph starting
"Also note the markers cannot open...") and update only that word to correct the
grammar.

Comment thread src/bun.js/api/MarkdownObject.zig
Comment thread src/bun.js/api/MarkdownObject.zig Outdated
Comment thread src/bundler/ParseTask.zig
Comment on lines +371 to +385
.md => {
const html = bun.md.renderToHtml(source.contents, allocator) catch {
log.addError(
source,
Logger.Loc.Empty,
"Failed to render markdown to HTML",
) catch |err| bun.handleOom(err);
return error.ParserError;
};
const root = Expr.init(E.String, E.String{
.data = html,
}, Logger.Loc{ .start = 0 });
var ast = JSAst.init((try js_parser.newLazyExportAST(allocator, transpiler.options.define, opts, log, root, source, "")).?);
ast.addUrlForCss(allocator, source, "text/html", null);
return ast;

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

⚠️ Potential issue | 🟠 Major

🧩 Analysis chain

🏁 Script executed:

#!/bin/bash
# Check ParseTask.zig around lines 371-385 to see the actual code
sed -n '365,390p' src/bundler/ParseTask.zig

Repository: oven-sh/bun

Length of output: 1227


🏁 Script executed:

#!/bin/bash
# Find the renderToHtml function definition
rg -n "fn renderToHtml" -g '*.zig' -A 3

Repository: oven-sh/bun

Length of output: 852


🏁 Script executed:

#!/bin/bash
# Search for renderToHtml calls in the codebase to see error handling patterns
rg -n "renderToHtml" -g '*.zig' -B 2 -A 2

Repository: oven-sh/bun

Length of output: 1998


Propagate markdown rendering failures instead of converting to ParserError.

The current code catches renderToHtml errors and converts them to error.ParserError, which masks OutOfMemory failures and prevents proper error propagation. Since HTML rendering is critical to the parse result, allocation failures must bubble up.

🔧 Proposed fix
-            const html = bun.md.renderToHtml(source.contents, allocator) catch {
-                log.addError(
-                    source,
-                    Logger.Loc.Empty,
-                    "Failed to render markdown to HTML",
-                ) catch |err| bun.handleOom(err);
-                return error.ParserError;
+            const html = bun.md.renderToHtml(source.contents, allocator) catch |err| {
+                log.addError(
+                    source,
+                    Logger.Loc.Empty,
+                    "Failed to render markdown to HTML",
+                ) catch |e| bun.handleOom(e);
+                return err;
🤖 Prompt for AI Agents
In `@src/bundler/ParseTask.zig` around lines 371 - 385, The markdown-to-HTML error
should not be converted into error.ParserError; change the catch on
bun.md.renderToHtml in the .md arm to a catch |err| block that logs the failure
(using log.addError with the original err context), calls bun.handleOom(err) to
let OOMs be handled, and then returns err to propagate the real error to the
caller instead of returning error.ParserError; update the code around
bun.md.renderToHtml, the catch block, and the surrounding control flow in
ParseTask.zig so the original error is returned.

Comment thread src/md/autolinks.zig
Comment thread src/md/blocks.zig
Comment thread src/transpiler.zig
Comment on lines +1358 to +1390
.md => {
const html = bun.md.renderToHtml(source.contents, allocator) catch {
transpiler.log.addErrorFmt(
null,
logger.Loc.Empty,
transpiler.allocator,
"Failed to render markdown to HTML",
.{},
) catch {};
return null;
};
const expr = js_ast.Expr.init(js_ast.E.String, js_ast.E.String{
.data = html,
}, logger.Loc.Empty);
const stmt = js_ast.Stmt.alloc(js_ast.S.ExportDefault, js_ast.S.ExportDefault{
.value = js_ast.StmtOrExpr{ .expr = expr },
.default_name = js_ast.LocRef{
.loc = logger.Loc{},
.ref = Ref.None,
},
}, logger.Loc{ .start = 0 });
var stmts = allocator.alloc(js_ast.Stmt, 1) catch unreachable;
stmts[0] = stmt;
var parts = allocator.alloc(js_ast.Part, 1) catch unreachable;
parts[0] = js_ast.Part{ .stmts = stmts };

return ParseResult{
.ast = js_ast.Ast.fromParts(parts),
.source = source.*,
.loader = loader,
.input_fd = input_fd,
};
},

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

⚠️ Potential issue | 🟡 Minor

Consider propagating the actual error from renderToHtml.

The error handling catches failures from renderToHtml but logs a generic message. The actual error from the Markdown renderer could help users diagnose issues (e.g., malformed input, allocation failure).

Suggested improvement
             .md => {
-                const html = bun.md.renderToHtml(source.contents, allocator) catch {
+                const html = bun.md.renderToHtml(source.contents, allocator) catch |err| {
                     transpiler.log.addErrorFmt(
                         null,
                         logger.Loc.Empty,
                         transpiler.allocator,
-                        "Failed to render markdown to HTML",
-                        .{},
+                        "Failed to render markdown to HTML: {s}",
+                        .{`@errorName`(err)},
                     ) catch {};
                     return null;
                 };
🤖 Prompt for AI Agents
In `@src/transpiler.zig` around lines 1358 - 1390, The catch on
bun.md.renderToHtml currently drops the renderer error; change the catch to
capture the error (e.g., catch |err| ...) and pass its message/details into
transpiler.log.addErrorFmt so the log includes the actual renderer error; update
the transpiler.log.addErrorFmt call used in the .md case (the block that
constructs js_ast.Expr and returns ParseResult) to include the captured err
information in the formatted message while preserving the existing context and
return null behavior.

Comment thread test/js/bun/md/gfm-compat.test.ts
Comment on lines +234 to +241
for (const { name, file } of specFiles) {
const specPath = join(SPEC_DIR, file);
let examples: SpecExample[];
try {
examples = parseSpecFile(specPath);
} catch {
continue;
}

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

⚠️ Potential issue | 🟡 Minor

Avoid silently skipping missing spec files.

catch { continue; } hides missing spec files or parse errors and drops whole suites. It’s safer to fail fast (or explicitly skip with a reason) so coverage doesn’t silently disappear.

🔧 Proposed fix
-  let examples: SpecExample[];
-  try {
-    examples = parseSpecFile(specPath);
-  } catch {
-    continue;
-  }
+  const examples = parseSpecFile(specPath);
📝 Committable suggestion

‼️ IMPORTANT
Carefully review the code before committing. Ensure that it accurately replaces the highlighted code, contains no missing lines, and has no issues with indentation. Thoroughly test & benchmark the code to ensure it meets the requirements.

Suggested change
for (const { name, file } of specFiles) {
const specPath = join(SPEC_DIR, file);
let examples: SpecExample[];
try {
examples = parseSpecFile(specPath);
} catch {
continue;
}
for (const { name, file } of specFiles) {
const specPath = join(SPEC_DIR, file);
const examples = parseSpecFile(specPath);
🤖 Prompt for AI Agents
In `@test/js/bun/md/md-spec.test.ts` around lines 234 - 241, The loop over
specFiles silently swallows parse errors by using catch { continue; } which
hides missing or invalid spec files; change the catch to capture the error
(e.g., catch (err)) and either rethrow a new Error that includes the specPath
and original error, or explicitly log a clear skip reason before continuing so
parseSpecFile failures are visible; update the block around
parseSpecFile(specPath) and the examples variable handling to include the error
variable and propagate or report it rather than silently continuing.

Comment thread test/js/bun/md/spec-permissive-autolinks.txt
Comment thread test/js/bun/md/spec-strikethrough.txt

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 1

🤖 Fix all issues with AI agents
In `@src/bun.js/api/MarkdownObject.zig`:
- Line 129: Remove the duplicated constant declaration BLOCK_FENCED_CODE from
the JsCallbackRenderer scope and replace all references to it with the canonical
md.types.BLOCK_FENCED_CODE; i.e., delete the local const BLOCK_FENCED_CODE =
0x10 in src/bun.js/api/MarkdownObject.zig and update uses inside the
JsCallbackRenderer struct and its methods to reference
md.types.BLOCK_FENCED_CODE (or import that symbol explicitly) so the flag value
remains authoritative from the md module.
♻️ Duplicate comments (4)
test/js/bun/md/gfm-compat.test.ts (1)

19-26: Enable tag_filter in renderGFM for tagfilter tests.

The renderGFM helper enables several GFM extensions but omits tag_filter. The tagfilter tests (lines 356-445) expect disallowed HTML tags to be filtered, which requires this option to be explicitly enabled.

🔧 Proposed fix
 function renderGFM(md: string): string {
   return render(md, {
     tables: true,
     strikethrough: true,
     tasklists: true,
     permissive_autolinks: true,
+    tag_filter: true,
   });
 }
test/js/bun/md/md-spec.test.ts (1)

234-241: Avoid silently skipping missing spec files.

The catch { continue; } pattern hides missing spec files or parse errors, which could cause test coverage to silently disappear. Consider failing fast or logging a clear skip reason.

🔧 Proposed fix
-  let examples: SpecExample[];
-  try {
-    examples = parseSpecFile(specPath);
-  } catch {
-    continue;
-  }
+  const examples = parseSpecFile(specPath);
src/bun.js/api/MarkdownObject.zig (2)

121-160: The private field naming issue was already raised in a previous review.


196-200: The silent OOM handling issue was already raised in a previous review (also applies to lines 254 and 274).

Comment thread src/bun.js/api/MarkdownObject.zig Outdated
Jarred-Sumner and others added 2 commits January 26, 2026 02:04
Parser fixes for GFM behavior:
- Tables can interrupt paragraphs (split paragraph, last line becomes header)
- Pipes inside code spans are cell delimiters (removed code span skip in table rows)
- Table header/delimiter column count validation
- Autolink URL paths allow ~, *, +, % characters
- Autolink post-processing: trim trailing unbalanced ) and entity-like suffixes
- Tag filter tracks raw depth for disallowed HTML tags in inline context
- Table \| escape replacement before inline processing

Ban-words compliance:
- Replace .jsBoolean(true/false) with .true/.false
- Replace .arguments_old() with .argumentsAsArray()
- Replace std.unicode.utf8Encode with inline encodeUtf8 helper
- Remove (Bun as any) casts from test files

Co-Authored-By: Claude <noreply@anthropic.com>
- Add Bun.markdown namespace types to bun.d.ts (Options, Callbacks, html, render)
- Add docs/runtime/markdown.mdx with API reference and examples
- Add markdown to docs.json navigation under Utilities
- Add Bun.markdown to bun-apis.mdx listing

Co-Authored-By: Claude <noreply@anthropic.com>
@Jarred-Sumner
Jarred-Sumner requested a review from alii as a code owner January 26, 2026 01:04

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 7

🤖 Fix all issues with AI agents
In `@docs/runtime/markdown.mdx`:
- Around line 192-203: The example shows the same key name "strikethrough" used
both as a callback and as a parser option in the Bun.markdown.render call, which
is confusing; update the example to avoid the name collision by renaming the
callback (e.g., "strikethroughCallback") or the option in the sample object, or
add a short clarifying sentence after the snippet stating that callbacks and
parser options share a namespace and that the last key wins so you should use
either the callback OR the option, not both; reference the Bun.markdown.render
invocation and the "strikethrough" callback/option to locate the change.

In `@src/md/html_renderer.zig`:
- Around line 377-393: In HtmlRenderer.updateTagFilterRawDepth, remove the
redundant length check inside the opening-tag branch (the `content.len < 2` in
the self-closing detection) because the function already returned for
content.len < 2 at the top; update the condition that detects NOT self-closing
to only inspect the last two bytes (e.g., check that content[content.len - 2] !=
'/' or content[content.len - 1] != '>') so the self-closing detection is correct
and simpler.

In `@src/md/line_analysis.zig`:
- Around line 59-60: The function isOpeningCodeFence reads self.text[off]
without verifying off is within bounds; add a defensive bounds check at the
start of isOpeningCodeFence (e.g., ensure off < self.text.len()) and return a
non-fence result immediately if out of range, then proceed to access
self.text[off] and subsequent indices only after verifying their bounds;
reference the isOpeningCodeFence function, the off parameter, and self.text
accesses when implementing this guard.

In `@src/md/render_blocks.zig`:
- Around line 104-117: The current code returns immediately on allocation
failure for buf.ensureTotalCapacity causing cell content to be skipped; change
the error handling for ensureTotalCapacity so you either propagate the
allocation error to the caller (returning the error) or set a local fallback
flag and continue processing without unescaping pipes (using original
cell_content) instead of an early return; update the handling where
ensureTotalCapacity is called (the buf.ensureTotalCapacity(self.allocator,
cell_content.len) invocation) and ensure subsequent logic around buf,
appendAssumeCapacity, and defer buf.deinit(self.allocator) respects the chosen
error path.

In `@test/js/bun/md/gfm-compat.test.ts`:
- Around line 13-17: The module-level const md = Bun.markdown is being shadowed
by test callback parameters named md; rename one of them for clarity—e.g., keep
const md = Bun.markdown but rename the test parameter(s) from md to markdown or
inputMd, and update any references in those tests (tests that call render or
assert against the parameter) so there is no shadowing; alternatively rename the
module-level const to markdownRenderer and keep test parameters as-is, ensuring
all usages of render and Bun.markdown reference the renamed identifier.

In `@test/js/bun/md/md-spec.test.ts`:
- Around line 15-69: The parseSpecFile function fails on CRLF files because
split("\n") leaves trailing "\r" characters which break sentinel checks like
lines[i] === "."; fix by normalizing line endings before splitting (e.g.,
replace all "\r\n" or "\r" with "\n") or by trimming "\r" from each line when
reading content; update parseSpecFile (references: parseSpecFile, readFileSync,
the split("\n") call and the "." sentinel checks) to perform this normalization
so the "." and fence comparisons work on Windows.

In `@test/js/bun/md/spec-tables.txt`:
- Around line 231-232: The sentence in spec-tables.txt ("Contents of each cell
is parsed as an inline text which may contents any") has a typo — replace
"contents" with "contain" so it reads "Contents of each cell is parsed as an
inline text which may contain any", updating the documentation text accordingly.
♻️ Duplicate comments (13)
test/js/bun/md/md-spec.test.ts (1)

234-241: Don’t silently skip spec parse errors.

Line 237–240 swallows missing/invalid spec files, which can hide coverage regressions. Fail fast (or log explicit skips). This was raised previously.

🔧 Proposed fix
-  let examples: SpecExample[];
-  try {
-    examples = parseSpecFile(specPath);
-  } catch {
-    continue;
-  }
+  const examples = parseSpecFile(specPath);
src/bun.js/api/MarkdownObject.zig (4)

112-151: Use # prefix for private struct fields.

Per coding guidelines, private fields in Zig structs should use the # prefix. The JsCallbackRenderer, Callbacks, and StackEntry structs have fields that are module-private.

As per coding guidelines, use #-prefixed field names for private structs.


120-120: Remove the duplicated BLOCK_FENCED_CODE constant.

This constant duplicates md.types.BLOCK_FENCED_CODE. Reference the canonical definition from the md module to avoid maintenance risk if the value changes.


187-191: Don't silently drop content on allocator failure.

appendToTop swallows OOM with catch {}, which can silently truncate output. Set an error flag on allocation failure.


241-246: Propagate allocation errors in VTable implementations.

The enterBlockImpl and enterSpanImpl functions swallow OOM on stack.append. This should set has_js_error (or a dedicated OOM flag) to avoid silent stack desynchronization.

Also applies to: 262-266

src/md/autolinks.zig (1)

136-146: Guard against out-of-range pos.

findPermissiveAutolink accesses content[pos] at line 139 without a bounds check. Since this is a pub function, add an early guard.

🐛 Suggested fix
 pub fn findPermissiveAutolink(content: []const u8, pos: usize, allow_emph: bool) AutolinkResult {
+    if (pos >= content.len) return .{ .found = false, .beg = 0, .end = 0 };
     const c = content[pos];
src/md/blocks.zig (2)

305-312: Use cont_result.off in the paragraph-interrupt bounds check.

The guard checks off >= self.size but then calls self.ch(cont_result.off). These are different values—cont_result.off is the offset after the container mark. The bounds check should use the same variable being accessed.

Suggested fix
-                    if ((off >= self.size or helpers.isNewline(self.ch(cont_result.off))) and container.ch != '>') {
+                    if ((cont_result.off >= self.size or helpers.isNewline(self.ch(cont_result.off))) and container.ch != '>') {
                         // Blank after list mark can't interrupt paragraph

763-767: Propagate OOM when appending block lines.

endCurrentBlock swallows allocation failure with catch {} on line 765, which can corrupt output silently.

🐛 Suggested fix
-        self.block_bytes.appendSlice(self.allocator, line_bytes) catch {};
+        try self.block_bytes.appendSlice(self.allocator, line_bytes);
src/md/line_analysis.zig (3)

1-3: Missing bounds check before array access.

isSetextUnderline accesses self.text[off] at line 2 without verifying off < self.size. Add a defensive guard.

🛡️ Suggested defensive check
 pub fn isSetextUnderline(self: *const Parser, off: OFF) struct { is_setext: bool, level: u32 } {
+    if (off >= self.size) return .{ .is_setext = false, .level = 0 };
     const c = self.text[off];

19-21: Same bounds check concern as isSetextUnderline.

isHrLine accesses self.text[off] at line 20 without bounds verification.

🛡️ Suggested defensive check
 pub fn isHrLine(self: *const Parser, off: OFF) bool {
+    if (off >= self.size) return false;
     const c = self.text[off];

311-374: Consider renaming isTableUnderline to reflect its side effects.

This function mutates self.table_alignments and self.table_col_count (lines 341-349, 372-373), which is unexpected for an is* predicate. Consider renaming to parseTableUnderline or documenting the side effects.

src/md/parser.zig (2)

52-53: Table alignment buffer size mismatch with TABLE_MAXCOLCOUNT.

table_alignments is sized to 64 but types.TABLE_MAXCOLCOUNT is 128. This can cause out-of-bounds writes for tables with more than 64 columns. Size the buffer to match the declared maximum.

🛠️ Suggested fix
-    table_col_count: u32 = 0,
-    table_alignments: [64]Align = [_]Align{.default} ** 64,
+    table_col_count: u32 = 0,
+    table_alignments: [types.TABLE_MAXCOLCOUNT]Align = [_]Align{.default} ** types.TABLE_MAXCOLCOUNT,

224-236: Use errdefer to prevent HtmlRenderer memory leak on error paths.

If processDoc() or toOwnedSlice() fails, the HtmlRenderer's internal buffer is leaked. The past review correctly identified this issue but suggested defer, which would incorrectly free the buffer before returning on success. Use errdefer to clean up only on error paths while allowing toOwnedSlice() to transfer ownership on success.

🛠️ Suggested fix
 pub fn renderToHtml(text: []const u8, allocator: Allocator, flags: Flags, tag_filter: bool) error{OutOfMemory}![]u8 {
     // Skip UTF-8 BOM
     const input = helpers.skipUtf8Bom(text);

     var html_renderer = HtmlRenderer.init(allocator, input, tag_filter);
+    errdefer html_renderer.deinit();

     var parser = Parser.init(allocator, input, flags, html_renderer.renderer());
     defer parser.deinit();

     try parser.processDoc();

     return html_renderer.toOwnedSlice();
 }

Comment thread docs/runtime/markdown.mdx
Comment thread src/md/html_renderer.zig
Comment thread src/md/line_analysis.zig
Comment thread src/md/render_blocks.zig Outdated
Comment thread test/js/bun/md/gfm-compat.test.ts Outdated
Comment thread test/js/bun/md/md-spec.test.ts
Comment on lines +231 to +232
Contents of each cell is parsed as an inline text which may contents any
inline Markdown spans like emphasis, strong emphasis, links etc.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

⚠️ Potential issue | 🟡 Minor

Minor typo in documentation.

"Contents of each cell is parsed as an inline text which may contents any" should read "...which may contain any".

🤖 Prompt for AI Agents
In `@test/js/bun/md/spec-tables.txt` around lines 231 - 232, The sentence in
spec-tables.txt ("Contents of each cell is parsed as an inline text which may
contents any") has a typo — replace "contents" with "contain" so it reads
"Contents of each cell is parsed as an inline text which may contain any",
updating the documentation text accordingly.

- Replace hand-rolled asciiCaseEql/matchTagNameCI with bun.strings APIs
- Extract shared parseEntityCodepoint and findEntity into helpers
- Convert found:bool result structs to Zig ?T optionals throughout
- Replace C-style while counting loops with for(0..n) ranges
- Convert if-chains to switch statements for boundary checks and label normalization
- Use std.mem.indexOfNone for whitespace-only checks

Co-Authored-By: Claude <noreply@anthropic.com>
Comment thread src/md/helpers.zig Outdated

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Can you hoist this up and make sure it's only referenced via &[_] so that we avoid this potentially living as stack memory?

Comment thread src/md/inlines.zig Outdated
pub fn findCodeSpanEnd(self: *const Parser, content: []const u8, start: usize, count: usize) ?usize {
_ = self;
var pos = start;
while (pos < content.len) {

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Can this use bun.strings.indexOfCharPos?

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 2

🤖 Fix all issues with AI agents
In `@src/md/ref_defs.zig`:
- Around line 9-61: lookupRefDef currently leaks the temporary normalized label
returned by normalizeLabel and normalizeLabel can leak its partially-built
ArrayList on OOM; fix by making normalizeLabel ensure the ArrayList is
deinitialized on all early-return/error paths (deinit the ArrayList before
returning raw on OOM) and change lookupRefDef to capture the normalized result
into a variable and defer its deinitialization (call the corresponding
ArrayList/allocator cleanup) before iterating; reference normalizeLabel and
lookupRefDef and ensure every path that returns from normalizeLabel or
lookupRefDef cleans up allocator-backed buffers via defer.

In `@src/md/render_blocks.zig`:
- Line 95: The literal 64 in the table alignment logic should be extracted to a
named constant for clarity and reuse: define a constant like MAX_TABLE_COLUMNS
(e.g., in the types or a common constants module) and replace the magic number
in the expression that computes align_data (the line using cell_index < 64 and
self.table_alignments) with that constant; update any other places that rely on
the same limit (e.g., accesses of self.table_alignments) to use the new
MAX_TABLE_COLUMNS to ensure consistency.
♻️ Duplicate comments (2)
src/md/autolinks.zig (1)

137-141: Add a bounds guard before indexing content[pos].

findPermissiveAutolink reads content[pos] without ensuring pos < content.len. Since this is public, an external caller can trigger OOB.

🐛 Proposed fix
 pub fn findPermissiveAutolink(content: []const u8, pos: usize, allow_emph: bool) AutolinkResult {
+    if (pos >= content.len) return null;
     const c = content[pos];
src/md/render_blocks.zig (1)

102-121: Allocation failure leaves block state unbalanced.

The enterBlock at line 96 is not matched by leaveBlock when ensureTotalCapacity fails at line 105. This leaves the renderer in an inconsistent state. The fallback should process the cell without pipe unescaping rather than returning early.

🔧 Suggested fix
             if (std.mem.indexOf(u8, cell_content, "\\|") != null) {
                 var buf: std.ArrayListUnmanaged(u8) = .{};
                 defer buf.deinit(self.allocator);
-                buf.ensureTotalCapacity(self.allocator, cell_content.len) catch return;
+                buf.ensureTotalCapacity(self.allocator, cell_content.len) catch {
+                    // Fallback: process without pipe unescaping on OOM
+                    self.processInlineContent(cell_content, vline.beg + `@as`(OFF, `@intCast`(cell_beg)));
+                    self.leaveBlock(cell_type, 0);
+                    cell_index += 1;
+                    if (end < row_text.len) {
+                        start = end + 1;
+                    } else {
+                        break;
+                    }
+                    continue;
+                };

Comment thread src/md/ref_defs.zig
Comment on lines +9 to +61
pub fn normalizeLabel(self: *Parser, raw: []const u8) []const u8 {
// Collapse whitespace and apply Unicode case folding (per CommonMark §6.7)
var result = std.ArrayListUnmanaged(u8){};
var in_ws = true; // skip leading whitespace
var i: usize = 0;
while (i < raw.len) {
const c = raw[i];
switch (c) {
' ', '\t', '\n', '\r' => {
if (!in_ws and result.items.len > 0) {
result.append(self.allocator, ' ') catch return raw;
in_ws = true;
}
i += 1;
},
0x80...0xFF => {
// Multi-byte UTF-8: decode, case fold, re-encode
const decoded = helpers.decodeUtf8(raw, i);
const fold = unicode.caseFold(decoded.codepoint);
var j: u2 = 0;
while (j < fold.n_codepoints) : (j += 1) {
var buf: [4]u8 = undefined;
const len = helpers.encodeUtf8(fold.codepoints[j], &buf);
if (len > 0) {
result.appendSlice(self.allocator, buf[0..len]) catch return raw;
}
}
in_ws = false;
i += @as(usize, decoded.len);
},
else => {
// ASCII: simple toLower
result.append(self.allocator, std.ascii.toLower(c)) catch return raw;
in_ws = false;
i += 1;
},
}
}
// Strip trailing space
if (result.items.len > 0 and result.items[result.items.len - 1] == ' ') {
result.items.len -= 1;
}
return result.items;
}

/// Look up a reference definition by label (case-insensitive, whitespace-normalized).
pub fn lookupRefDef(self: *Parser, raw_label: []const u8) ?RefDef {
if (raw_label.len == 0) return null;
const normalized = self.normalizeLabel(raw_label);
if (normalized.len == 0) return null; // whitespace-only labels are invalid
for (self.ref_defs.items) |rd| {
if (std.mem.eql(u8, rd.label, normalized)) return rd;
}

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

⚠️ Potential issue | 🟠 Major

Free temporary normalized labels to avoid per‑lookup leaks.

lookupRefDef allocates a fresh normalized label on every call and never releases it, so repeated lookups grow memory. Also, early OOM returns in normalizeLabel can leak partially-built buffers. Use a defer to free temporary normalized labels in lookupRefDef and deinit the list in OOM paths. As per coding guidelines, ensure allocator-backed buffers are cleaned up with defer.

🐛 Proposed fix
 pub fn normalizeLabel(self: *Parser, raw: []const u8) []const u8 {
     // Collapse whitespace and apply Unicode case folding (per CommonMark §6.7)
     var result = std.ArrayListUnmanaged(u8){};
     var in_ws = true; // skip leading whitespace
@@
         switch (c) {
             ' ', '\t', '\n', '\r' => {
                 if (!in_ws and result.items.len > 0) {
-                    result.append(self.allocator, ' ') catch return raw;
+                    result.append(self.allocator, ' ') catch {
+                        result.deinit(self.allocator);
+                        return raw;
+                    };
                     in_ws = true;
                 }
                 i += 1;
             },
             0x80...0xFF => {
@@
                     const len = helpers.encodeUtf8(fold.codepoints[j], &buf);
                     if (len > 0) {
-                        result.appendSlice(self.allocator, buf[0..len]) catch return raw;
+                        result.appendSlice(self.allocator, buf[0..len]) catch {
+                            result.deinit(self.allocator);
+                            return raw;
+                        };
                     }
                 }
                 in_ws = false;
                 i += `@as`(usize, decoded.len);
             },
             else => {
                 // ASCII: simple toLower
-                result.append(self.allocator, std.ascii.toLower(c)) catch return raw;
+                result.append(self.allocator, std.ascii.toLower(c)) catch {
+                    result.deinit(self.allocator);
+                    return raw;
+                };
                 in_ws = false;
                 i += 1;
             },
         }
     }
@@
 pub fn lookupRefDef(self: *Parser, raw_label: []const u8) ?RefDef {
     if (raw_label.len == 0) return null;
     const normalized = self.normalizeLabel(raw_label);
+    const owns_normalized = normalized.ptr != raw_label.ptr or normalized.len != raw_label.len;
+    defer if (owns_normalized) self.allocator.free(normalized);
     if (normalized.len == 0) return null; // whitespace-only labels are invalid
     for (self.ref_defs.items) |rd| {
         if (std.mem.eql(u8, rd.label, normalized)) return rd;
     }
     return null;
 }
📝 Committable suggestion

‼️ IMPORTANT
Carefully review the code before committing. Ensure that it accurately replaces the highlighted code, contains no missing lines, and has no issues with indentation. Thoroughly test & benchmark the code to ensure it meets the requirements.

Suggested change
pub fn normalizeLabel(self: *Parser, raw: []const u8) []const u8 {
// Collapse whitespace and apply Unicode case folding (per CommonMark §6.7)
var result = std.ArrayListUnmanaged(u8){};
var in_ws = true; // skip leading whitespace
var i: usize = 0;
while (i < raw.len) {
const c = raw[i];
switch (c) {
' ', '\t', '\n', '\r' => {
if (!in_ws and result.items.len > 0) {
result.append(self.allocator, ' ') catch return raw;
in_ws = true;
}
i += 1;
},
0x80...0xFF => {
// Multi-byte UTF-8: decode, case fold, re-encode
const decoded = helpers.decodeUtf8(raw, i);
const fold = unicode.caseFold(decoded.codepoint);
var j: u2 = 0;
while (j < fold.n_codepoints) : (j += 1) {
var buf: [4]u8 = undefined;
const len = helpers.encodeUtf8(fold.codepoints[j], &buf);
if (len > 0) {
result.appendSlice(self.allocator, buf[0..len]) catch return raw;
}
}
in_ws = false;
i += @as(usize, decoded.len);
},
else => {
// ASCII: simple toLower
result.append(self.allocator, std.ascii.toLower(c)) catch return raw;
in_ws = false;
i += 1;
},
}
}
// Strip trailing space
if (result.items.len > 0 and result.items[result.items.len - 1] == ' ') {
result.items.len -= 1;
}
return result.items;
}
/// Look up a reference definition by label (case-insensitive, whitespace-normalized).
pub fn lookupRefDef(self: *Parser, raw_label: []const u8) ?RefDef {
if (raw_label.len == 0) return null;
const normalized = self.normalizeLabel(raw_label);
if (normalized.len == 0) return null; // whitespace-only labels are invalid
for (self.ref_defs.items) |rd| {
if (std.mem.eql(u8, rd.label, normalized)) return rd;
}
pub fn normalizeLabel(self: *Parser, raw: []const u8) []const u8 {
// Collapse whitespace and apply Unicode case folding (per CommonMark §6.7)
var result = std.ArrayListUnmanaged(u8){};
var in_ws = true; // skip leading whitespace
var i: usize = 0;
while (i < raw.len) {
const c = raw[i];
switch (c) {
' ', '\t', '\n', '\r' => {
if (!in_ws and result.items.len > 0) {
result.append(self.allocator, ' ') catch {
result.deinit(self.allocator);
return raw;
};
in_ws = true;
}
i += 1;
},
0x80...0xFF => {
// Multi-byte UTF-8: decode, case fold, re-encode
const decoded = helpers.decodeUtf8(raw, i);
const fold = unicode.caseFold(decoded.codepoint);
var j: u2 = 0;
while (j < fold.n_codepoints) : (j += 1) {
var buf: [4]u8 = undefined;
const len = helpers.encodeUtf8(fold.codepoints[j], &buf);
if (len > 0) {
result.appendSlice(self.allocator, buf[0..len]) catch {
result.deinit(self.allocator);
return raw;
};
}
}
in_ws = false;
i += `@as`(usize, decoded.len);
},
else => {
// ASCII: simple toLower
result.append(self.allocator, std.ascii.toLower(c)) catch {
result.deinit(self.allocator);
return raw;
};
in_ws = false;
i += 1;
},
}
}
// Strip trailing space
if (result.items.len > 0 and result.items[result.items.len - 1] == ' ') {
result.items.len -= 1;
}
return result.items;
}
/// Look up a reference definition by label (case-insensitive, whitespace-normalized).
pub fn lookupRefDef(self: *Parser, raw_label: []const u8) ?RefDef {
if (raw_label.len == 0) return null;
const normalized = self.normalizeLabel(raw_label);
const owns_normalized = normalized.ptr != raw_label.ptr or normalized.len != raw_label.len;
defer if (owns_normalized) self.allocator.free(normalized);
if (normalized.len == 0) return null; // whitespace-only labels are invalid
for (self.ref_defs.items) |rd| {
if (std.mem.eql(u8, rd.label, normalized)) return rd;
}
return null;
}
🤖 Prompt for AI Agents
In `@src/md/ref_defs.zig` around lines 9 - 61, lookupRefDef currently leaks the
temporary normalized label returned by normalizeLabel and normalizeLabel can
leak its partially-built ArrayList on OOM; fix by making normalizeLabel ensure
the ArrayList is deinitialized on all early-return/error paths (deinit the
ArrayList before returning raw on OOM) and change lookupRefDef to capture the
normalized result into a variable and defer its deinitialization (call the
corresponding ArrayList/allocator cleanup) before iterating; reference
normalizeLabel and lookupRefDef and ensure every path that returns from
normalizeLabel or lookupRefDef cleans up allocator-backed buffers via defer.

Comment thread src/md/render_blocks.zig Outdated

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 7

🤖 Fix all issues with AI agents
In `@docs/runtime/markdown.mdx`:
- Around line 72-74: The description for Bun.markdown.render() incorrectly
implies it returns React elements; update the text for the
`Bun.markdown.render()` section to clarify that it concatenates string outputs
and returns a string—replace or remove the phrase "React elements" (suggested
replacement: "HTML/JSX strings" or simply "HTML strings") so the docs accurately
state the function returns a string.

In `@src/bun.js/api/MarkdownObject.zig`:
- Around line 84-90: The code returns error.JSError when
js_renderer.has_js_error is set but doesn't distinguish OOM from JS exceptions;
update the render-return paths (the block that checks js_renderer.has_js_error
and then calls js_renderer.getResult() / bun.String.createUTF8ForJS) to track
allocator OOM separately (e.g., set a separate flag when appendToTop/stack
append fails) and, before returning error.JSError, call throwOutOfMemory() if
that OOM flag is set; apply the same change to the other occurrences where
js_renderer.has_js_error is checked so any internal allocator failure triggers
throwOutOfMemory() rather than a plain error.JSError.

In `@src/md/blocks.zig`:
- Around line 770-839: The consumeRefDefsFromCurrentBlock function currently
swallows allocator failures (self.buffer.append(... ) catch {},
self.allocator.dupe(...) catch return, self.ref_defs.append(...) catch return)
which can lead to lost/misaligned data; change these to handle errors
explicitly: for buffer append failures (self.buffer.append and the second
appendSlice call) return an error or set a parser-level warning flag and stop
merging so parsing doesn't continue on a partially-built merged buffer; for
dupe/append failures (self.allocator.dupe for result.dest/result.title and
self.ref_defs.append) propagate the error (return) or record a non-fatal warning
and skip adding that single ref-def instead of silently continuing; update
callers of consumeRefDefsFromCurrentBlock or Parser to accept/observe the
returned error or flag so callers can log or handle allocation failures.

In `@src/md/inlines.zig`:
- Around line 24-46: The function processLeafBlock currently swallows allocator
failures via the catch {} on buffer.append and buffer.appendSlice which can
produce incomplete merged content; change processLeafBlock to propagate
allocation errors (e.g., return an error union like ?void or an error type)
instead of swallowing them, replace the catch {} with proper error propagation
from buffer.append and buffer.appendSlice, and update all callers to handle the
propagated error (including calls to processInlineContent) or alternatively
record a failure flag and early-return before calling processInlineContent;
reference the symbols processLeafBlock, buffer.append, buffer.appendSlice,
merged, and processInlineContent when making the changes.
- Around line 461-468: The append calls to self.emph_delims
(self.emph_delims.append(self.allocator, {...}) catch {}) silently ignore
allocation failures which drops delimiters and corrupts emphasis parsing;
replace the silent catch by handling the error—either propagate the allocator
error out of the surrounding function (use try/return err) or record a
persistent allocation-failure flag on the parser (e.g., self.alloc_failed =
true) and abort/skip further delimiter collection so callers can react; update
both occurrences (the append at run_start and the second append around lines
477-484) to use the chosen error-handling approach and ensure the caller checks
the propagated error or the failure flag.

In `@src/md/parser.zig`:
- Around line 5-62: The Parser struct exposes many internal fields (e.g.,
allocator, text, size, flags, renderer, image_nesting_level, link_nesting_level,
code_indent_offset, mark_char_map, marks, containers, block_bytes, buffer,
emph_delims, n_containers, current_block, current_block_lines, opener_stacks,
unresolved_link_head, table_cell_boundaries_head, html_block_type, fence_indent,
table_col_count, table_alignments, ref_defs,
last_line_has_list_loosening_effect, etc.); rename these fields to use the Zig
private-prefix (prepend `#` to each field name) and then update every usage
across the codebase (constructors/initializers, method implementations, tests,
and any module that referenced them) to the new `#`-prefixed names so the Parser
API remains public while its internal state is private. Ensure struct literals,
default values, and any pattern-matching or field access are updated
consistently and run the build to fix all references.

In `@test/js/bun/md/gfm-compat.test.ts`:
- Around line 354-356: Update the comment to reflect current behavior: replace
the outdated note about Bun needing a new tagfilter option with a statement that
renderGFM explicitly enables tag_filter, so the tests document GFM behavior with
tag_filter already enabled; reference renderGFM and the tag_filter option in the
updated comment near the existing NOTE block.
♻️ Duplicate comments (5)
src/md/line_analysis.zig (2)

1-3: Guard self.text[off] access in isSetextUnderline and isHrLine.
Both functions index with off before bounds checks. If callers ever pass off == self.size, this will panic.

🛡️ Suggested guards
 pub fn isSetextUnderline(self: *const Parser, off: OFF) struct { is_setext: bool, level: u32 } {
+    if (off >= self.size) return .{ .is_setext = false, .level = 0 };
     const c = self.text[off];
@@
 pub fn isHrLine(self: *const Parser, off: OFF) bool {
+    if (off >= self.size) return false;
     const c = self.text[off];

Also applies to: 19-21


312-374: isTableUnderline mutates parser state — rename or document.
This function updates self.table_alignments and self.table_col_count, but the is* name implies a pure predicate.

test/js/bun/md/md-spec.test.ts (1)

234-241: Don’t silently skip spec files on parse errors.
catch { continue; } hides missing/invalid spec data and drops coverage.

🛡️ Suggested fix
-  let examples: SpecExample[];
-  try {
-    examples = parseSpecFile(specPath);
-  } catch {
-    continue;
-  }
+  const examples = parseSpecFile(specPath);
src/bun.js/api/MarkdownObject.zig (1)

112-149: Use #‑prefixed private fields for module‑private structs.
JsCallbackRenderer, Callbacks, and StackEntry are module-private; please rename fields with # and update usages. As per coding guidelines.

♻️ Example adjustment
 const JsCallbackRenderer = struct {
-    globalObject: *jsc.JSGlobalObject,
-    allocator: std.mem.Allocator,
-    src_text: []const u8,
-    stack: std.ArrayListUnmanaged(StackEntry) = .{},
-    callbacks: Callbacks = .{},
-    has_js_error: bool = false,
+    `#globalObject`: *jsc.JSGlobalObject,
+    `#allocator`: std.mem.Allocator,
+    `#src_text`: []const u8,
+    `#stack`: std.ArrayListUnmanaged(StackEntry) = .{},
+    `#callbacks`: Callbacks = .{},
+    `#has_js_error`: bool = false,
 };
src/md/inlines.zig (1)

554-558: Wrapping subtraction trick needs clarifying comment.

The -%= 1 operation is used to undo the loop's upcoming increment so the same closer gets re-processed. While there's a comment, it could be clearer about the intent.

📝 Suggested clarification
                 // If closer still has remaining, re-process it (don't increment closer_idx)
                 if (delims[closer_idx].remaining > 0 and delims[closer_idx].can_close) {
-                    // Reset the while condition — we'll re-check this closer
-                    closer_idx -%= 1; // will be incremented by while loop
+                    // Decrement so the while loop's `: (closer_idx += 1)` brings us back
+                    // to this same index, allowing another matching attempt with the
+                    // remaining delimiter characters
+                    closer_idx -%= 1;
                 }

Comment thread docs/runtime/markdown.mdx
Comment on lines +72 to +74
## `Bun.markdown.render()`

Parse Markdown and render it using custom JavaScript callbacks. This gives you full control over the output format — you can generate HTML with custom classes, React elements, ANSI terminal output, or any other string format.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

⚠️ Potential issue | 🟡 Minor

Avoid implying render() returns React elements.
render() concatenates string outputs and returns a string; “React elements” suggests object returns. Consider wording like “HTML/JSX strings” or remove the React reference to avoid confusion.

📝 Suggested doc tweak
-Parse Markdown and render it using custom JavaScript callbacks. This gives you full control over the output format — you can generate HTML with custom classes, React elements, ANSI terminal output, or any other string format.
+Parse Markdown and render it using custom JavaScript callbacks. This gives you full control over the output format — you can generate HTML with custom classes, JSX/HTML strings, ANSI terminal output, or any other string format.
📝 Committable suggestion

‼️ IMPORTANT
Carefully review the code before committing. Ensure that it accurately replaces the highlighted code, contains no missing lines, and has no issues with indentation. Thoroughly test & benchmark the code to ensure it meets the requirements.

Suggested change
## `Bun.markdown.render()`
Parse Markdown and render it using custom JavaScript callbacks. This gives you full control over the output format — you can generate HTML with custom classes, React elements, ANSI terminal output, or any other string format.
## `Bun.markdown.render()`
Parse Markdown and render it using custom JavaScript callbacks. This gives you full control over the output format — you can generate HTML with custom classes, JSX/HTML strings, ANSI terminal output, or any other string format.
🤖 Prompt for AI Agents
In `@docs/runtime/markdown.mdx` around lines 72 - 74, The description for
Bun.markdown.render() incorrectly implies it returns React elements; update the
text for the `Bun.markdown.render()` section to clarify that it concatenates
string outputs and returns a string—replace or remove the phrase "React
elements" (suggested replacement: "HTML/JSX strings" or simply "HTML strings")
so the docs accurately state the function returns a string.

Comment thread src/bun.js/api/MarkdownObject.zig Outdated
Comment thread src/md/blocks.zig
Comment on lines +770 to +839
pub fn consumeRefDefsFromCurrentBlock(self: *Parser) void {
const items = self.current_block_lines.items;
if (items.len == 0) return;

// Merge lines into buffer for ref def parsing
self.buffer.clearRetainingCapacity();
for (items) |vline| {
if (vline.beg > vline.end or vline.end > self.size) continue;
if (self.buffer.items.len > 0) {
self.buffer.append(self.allocator, '\n') catch {};
}
self.buffer.appendSlice(self.allocator, self.text[vline.beg..vline.end]) catch {};
}

const merged = self.buffer.items;
var pos: usize = 0;
var lines_consumed: u32 = 0;

while (pos < merged.len) {
const result = self.parseRefDef(merged, pos) orelse break;

const norm_label = self.normalizeLabel(result.label);
if (norm_label.len == 0) break;

// First definition wins
var already_exists = false;
for (self.ref_defs.items) |existing| {
if (std.mem.eql(u8, existing.label, norm_label)) {
already_exists = true;
break;
}
}
if (!already_exists) {
const dest_dupe = self.allocator.dupe(u8, result.dest) catch return;
const title_dupe = self.allocator.dupe(u8, result.title) catch return;
self.ref_defs.append(self.allocator, .{
.label = norm_label,
.dest = dest_dupe,
.title = title_dupe,
}) catch return;
}

var newlines: u32 = 0;
for (merged[pos..result.end_pos]) |mc| {
if (mc == '\n') newlines += 1;
}
if (result.end_pos >= merged.len and (result.end_pos == pos or merged[result.end_pos - 1] != '\n')) {
newlines += 1;
}
lines_consumed += newlines;
pos = result.end_pos;
}

if (lines_consumed > 0) {
if (self.current_block) |cb_off| {
var hdr = self.getBlockHeaderAt(cb_off);
if (lines_consumed >= hdr.n_lines) {
// All lines consumed
self.current_block_lines.clearRetainingCapacity();
hdr.n_lines = 0;
} else {
// Remove first lines_consumed lines
const remaining = items.len - lines_consumed;
std.mem.copyForwards(VerbatimLine, items[0..remaining], items[lines_consumed..]);
self.current_block_lines.shrinkRetainingCapacity(remaining);
hdr.n_lines -= lines_consumed;
}
}
}
}

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🧹 Nitpick | 🔵 Trivial

Silent allocation failures in consumeRefDefsFromCurrentBlock may cause data loss.

Lines 779, 781, and 803-809 use catch {} or catch return for allocations. While ref-def processing is supplementary metadata, silently dropping allocations here could cause:

  • Incomplete buffer merging (lines 779, 781)
  • Missing ref-def entries (lines 803-809)

Consider at minimum logging a warning or setting a flag when these allocations fail.

♻️ Suggested improvement
 pub fn consumeRefDefsFromCurrentBlock(self: *Parser) void {
     const items = self.current_block_lines.items;
     if (items.len == 0) return;

     // Merge lines into buffer for ref def parsing
     self.buffer.clearRetainingCapacity();
+    var merge_failed = false;
     for (items) |vline| {
         if (vline.beg > vline.end or vline.end > self.size) continue;
         if (self.buffer.items.len > 0) {
-            self.buffer.append(self.allocator, '\n') catch {};
+            self.buffer.append(self.allocator, '\n') catch {
+                merge_failed = true;
+                break;
+            };
         }
-        self.buffer.appendSlice(self.allocator, self.text[vline.beg..vline.end]) catch {};
+        self.buffer.appendSlice(self.allocator, self.text[vline.beg..vline.end]) catch {
+            merge_failed = true;
+            break;
+        };
     }
+    if (merge_failed) return; // Can't parse incomplete buffer
🤖 Prompt for AI Agents
In `@src/md/blocks.zig` around lines 770 - 839, The consumeRefDefsFromCurrentBlock
function currently swallows allocator failures (self.buffer.append(... ) catch
{}, self.allocator.dupe(...) catch return, self.ref_defs.append(...) catch
return) which can lead to lost/misaligned data; change these to handle errors
explicitly: for buffer append failures (self.buffer.append and the second
appendSlice call) return an error or set a parser-level warning flag and stop
merging so parsing doesn't continue on a partially-built merged buffer; for
dupe/append failures (self.allocator.dupe for result.dest/result.title and
self.ref_defs.append) propagate the error (return) or record a non-fatal warning
and skip adding that single ref-def instead of silently continuing; update
callers of consumeRefDefsFromCurrentBlock or Parser to accept/observe the
returned error or flag so callers can log or handle allocation failures.

Comment thread src/md/inlines.zig Outdated
Comment thread src/md/inlines.zig
Comment thread src/md/parser.zig
Comment on lines +5 to +62
allocator: Allocator,
text: []const u8,
size: OFF,
flags: Flags,

// Output
renderer: Renderer,
image_nesting_level: u32 = 0,
link_nesting_level: u32 = 0,

// Code indent offset: 4 normally, maxInt if no_indented_code_blocks
code_indent_offset: u32,
doc_ends_with_newline: bool,

// Mark character map
mark_char_map: [256]bool = [_]bool{false} ** 256,

// Dynamic arrays
marks: std.ArrayListUnmanaged(Mark) = .{},
containers: std.ArrayListUnmanaged(Container) = .{},
block_bytes: std.ArrayListAlignedUnmanaged(u8, .@"4") = .{},
buffer: std.ArrayListUnmanaged(u8) = .{},
emph_delims: std.ArrayListUnmanaged(EmphDelim) = .{},

// Number of active containers
n_containers: u32 = 0,

// Current block being built
current_block: ?usize = null,
current_block_lines: std.ArrayListUnmanaged(VerbatimLine) = .{},

// Opener stacks
opener_stacks: [types.NUM_OPENER_STACKS]types.OpenerStack =
[_]types.OpenerStack{.{}} ** types.NUM_OPENER_STACKS,

// Linked lists through marks
unresolved_link_head: i32 = -1,
unresolved_link_tail: i32 = -1,
table_cell_boundaries_head: i32 = -1,
table_cell_boundaries_tail: i32 = -1,

// HTML block tracking
html_block_type: u8 = 0,
// Fenced code block indent
fence_indent: u32 = 0,

// Table column alignments
table_col_count: u32 = 0,
table_alignments: [types.TABLE_MAXCOLCOUNT]Align = [_]Align{.default} ** types.TABLE_MAXCOLCOUNT,

// Ref defs
ref_defs: std.ArrayListUnmanaged(RefDef) = .{},

// State
last_line_has_list_loosening_effect: bool = false,
last_list_item_starts_with_two_blank_lines: bool = false,
max_ref_def_output: u64 = 0,

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🛠️ Refactor suggestion | 🟠 Major

Use # prefix for private Parser fields.

Parser is public but its state fields appear internal; per the Zig style in this repo, private fields should be prefixed with #. Consider renaming these fields and updating usages across modules to avoid exposing internal state.

♻️ Example (apply consistently)
-    allocator: Allocator,
-    text: []const u8,
+    `#allocator`: Allocator,
+    `#text`: []const u8,

As per coding guidelines, private fields in Zig structs should use the # prefix.

🤖 Prompt for AI Agents
In `@src/md/parser.zig` around lines 5 - 62, The Parser struct exposes many
internal fields (e.g., allocator, text, size, flags, renderer,
image_nesting_level, link_nesting_level, code_indent_offset, mark_char_map,
marks, containers, block_bytes, buffer, emph_delims, n_containers,
current_block, current_block_lines, opener_stacks, unresolved_link_head,
table_cell_boundaries_head, html_block_type, fence_indent, table_col_count,
table_alignments, ref_defs, last_line_has_list_loosening_effect, etc.); rename
these fields to use the Zig private-prefix (prepend `#` to each field name) and
then update every usage across the codebase (constructors/initializers, method
implementations, tests, and any module that referenced them) to the new
`#`-prefixed names so the Parser API remains public while its internal state is
private. Ensure struct literals, default values, and any pattern-matching or
field access are updated consistently and run the build to fix all references.

Comment thread test/js/bun/md/gfm-compat.test.ts Outdated
Jarred-Sumner and others added 3 commits January 26, 2026 05:22
Add two new renderer options for generating GitHub-compatible heading
anchors. heading_ids generates id="slug" attributes on heading tags,
and autolink_headings wraps heading content in <a href="#slug"> anchors.

Slug algorithm: lowercase, strip non-alphanumeric (keep hyphens/spaces),
collapse consecutive hyphens, deduplicate with -1, -2 suffixes.

Co-Authored-By: Claude <noreply@anthropic.com>
Accept camelCase option names (headingIds, tagFilter, permissiveAutolinks,
etc.) with snake_case fallback for backward compatibility. Update
TypeScript types and tests to use camelCase.

Co-Authored-By: Claude <noreply@anthropic.com>

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 23

🤖 Fix all issues with AI agents
In `@bench/json5/json5.mjs`:
- Around line 12-57: The benchmark data is non-deterministic because
Math.random() is used in generateLargeJson5 and when building largeObject;
replace those Math.random() calls with a seeded PRNG so runs are reproducible:
add a small deterministic generator (e.g., mulberry32/xorshift) initialized with
a fixed seed and call it instead of Math.random() inside generateLargeJson5 (for
value and hex generation) and inside the largeObject Array.from mapping (for
value), ensuring hex formatting still uses the PRNG output and all places that
used Math.random() are switched to the seeded generator (refer to function
generateLargeJson5 and the largeObject constant).

In `@docs/runtime/file-types.mdx`:
- Line 8: The docs list of supported file extensions is missing `.md`, which is
misleading because the `.md` loader exists; update the extensions list in the
file-types overview by adding `.md` to the existing inline list (the string
containing `.js .cjs .mjs ... .sh`) so the markdown loader is documented
alongside the other extensions.

In `@packages/bun-types/bun.d.ts`:
- Around line 1149-1192: RenderCallbacks is missing handlers for parsed math and
underline spans, so when latexMath/underline are enabled render() silently drops
those nodes; add two callbacks to the RenderCallbacks interface—math (e.g.
math?: (children: string, meta?: MathMeta) => string | null | undefined) and
underline (e.g. underline?: (children: string) => string | null | undefined)—so
getSpanCallback()/render() can map those span types instead of returning .zero;
update types referenced (MathMeta) or reuse existing meta shapes as appropriate
and ensure RenderCallbacks includes these new symbols.

In `@src/api/schema.zig`:
- Around line 345-348: The C/C++ binding enum BunLoaderType in
headers-handwritten.h is missing several members and has numeric mismatches with
the Zig Loader enum in src/api/schema.zig; add constants for base64, dataurl,
text, bunsh, sqlite, sqlite_embedded, html, and json5 with the exact integer
values that match the Zig Loader sequence (fill in values 12–18 for the missing
ones and add json5 at its correct value), and update the md constant to use the
same numeric value as the Zig enum (ensure all BunLoaderType numeric assignments
mirror the Zig Loader order so there are no conflicts between md and json5).

In `@src/bun.js/bindings/napi.cpp`:
- Around line 2106-2109: The ArrayBuffer creation using
ArrayBuffer::createFromBytes with createSharedTask can throw, but the code does
not guard with NAPI_RETURN_IF_EXCEPTION(env), so if creation fails the shared
finalizer may run and call finalize_cb on caller-owned memory; immediately after
the ArrayBuffer::createFromBytes(...) call (the line constructing arrayBuffer)
add NAPI_RETURN_IF_EXCEPTION(env) to return on exception and prevent scheduling
the createSharedTask finalizer; reference symbols: ArrayBuffer::createFromBytes,
createSharedTask, finalize_cb, finalize_hint, env.

In `@src/interchange/yaml.zig`:
- Around line 404-408: The .stack_overflow branch in addToLog currently returns
error.StackOverflow without writing any diagnostic entries; modify the
.stack_overflow arm in pub fn addToLog(this: *const Error, source: *const
logger.Source, log: *logger.Log) to emit the same kind of diagnostic/log entry
used by other branches (e.g., unexpected_eof) before returning
error.StackOverflow so that temp_log is populated for the bundler path and
YAML.parse/SyntaxError fallback; locate the .stack_overflow case in addToLog and
add a call to the logger routines that produce the diagnostic (using the
provided source and log) prior to returning error.StackOverflow.

In `@src/md/helpers.zig`:
- Around line 447-460: trackText currently swallows allocation failures via
multiple "appendSlice(... ) catch {}" calls which can silently drop heading
text; change trackText (in HeadingIdTracker) to surface or handle OOMs instead
of ignoring them: either make trackText return an error (e.g., pub fn
trackText(...) !void) and replace each "appendSlice(... ) catch {}" with "try
self.text_buf.appendSlice(allocator, ...)" so allocation errors propagate, or
centrally handle the error (e.g., catch |err| { log the allocator error and set
self.in_heading = false } ) so you do not silently lose heading content; update
all sites calling trackText accordingly and include the same treatment for the
decodeEntityToUtf8 branch and references to self.text_buf.appendSlice.

In `@src/options.zig`:
- Around line 782-783: Update the invalid-loader error string passed to
global.throwInvalidArguments so it includes "json5" among the allowed loader
names; specifically modify the string currently returned (the one that lists
"js, jsx, tsx, ts, css, file, toml, yaml, wasm, bunsh, json, or md") to also
include "json5" (e.g., "... json5, or md") so users see the correct set of valid
loaders when the check in this code path triggers.

In `@test/js/bun/jsc-stress/fixtures/dfg-put-by-val-direct-with-edge-numbers.js`:
- Line 94: The inner for-loop reuses the outer-scope variable name `i`, which is
confusing and can trigger linters; change the loop variable in the inner loop
that iterates "for (var i = 0; i < testLoopCount; ++i)" to a distinct name such
as `j` (or another unused identifier) so the loop uses `for (var j = 0; j <
testLoopCount; ++j)` and update any references inside that loop body accordingly
(search for uses of `i` within the inner loop scope to rename).

In
`@test/js/bun/jsc-stress/fixtures/dfg-try-catch-wrong-value-recovery-on-ic-miss.js`:
- Around line 20-25: The object literal o2 defines a data property f: 500 and
immediately declares a getter get f() that overwrites it, making the 500 value
unreachable; either remove the redundant data property f: 500 from o2 or, if the
duplicate was intentional to shape the object for IC/stress testing, keep both
but add an inline comment next to o2 (or above it) explaining the intent (e.g.,
"data property present to force shape, then replaced by getter for IC recovery
tests") so future readers understand why f appears twice.

In `@test/js/bun/jsc-stress/fixtures/ftl-call-exception.js`:
- Around line 20-47: The code redeclares result with var and reassigns the
function-declared symbol bar; change the top-level function declaration bar to a
function expression assigned to a const (e.g., const bar = function(...) { ... }
or const bar = (..) => ... ) so it can be reassigned later, and replace both var
result declarations in the warm-up loop and after the throw with block-scoped
let (or const where appropriate) to avoid redeclaration errors; update
references to result and bar in foo(...) calls accordingly.

In `@test/js/bun/jsc-stress/fixtures/ftl-get-by-id-getter-exception.js`:
- Around line 21-28: The loop redeclares function-scoped vars (o, result) with
var each iteration which violates noRedeclare; fix by declaring these bindings
once and reusing them (e.g., move declaration of o and result outside the for
loop or switch to block-scoped let for o and result inside the loop) so that the
code using foo(o) and testLoopCount keeps the same behavior without var
redeclarations.

In `@test/js/bun/jsc-stress/fixtures/ftl-get-by-id-slow-exception.js`:
- Around line 21-25: The loop re-declares function-scoped vars (var o, var
result) causing noRedeclare lint failures; fix by declaring the bindings once
before the loop and reusing them inside (e.g., replace "var o; o = {…}; var
result = foo(o);" with a single pre-loop declaration like "let o, result;" or
"var o, result;" hoisted once and then assign within the loop), update the same
pattern in the other occurrence that mirrors lines 41-45, and ensure calls to
foo(o) still assign into the reused result variable.

In `@test/js/bun/jsc-stress/fixtures/ftl-string-equality.js`:
- Around line 9-13: The loops currently use "var i" which redeclares i in the
same scope; change each loop header that uses "var i" (e.g., the for loop
initializing array elements: for (var i = 0; ...), and the other loops later in
the file referenced at Lines 26 and 29) to use "let i" instead so each loop gets
a block-scoped iterator and avoids Biome noRedeclare errors.

In `@test/js/bun/jsc-stress/fixtures/ftl-try-catch-varargs-call-throws.js`:
- Around line 29-33: Add an explicit assertion that the final call to foo(f,
[10, 20, 30]) throws once flag is set so the exceptional path is actually
tested; locate the loop and final call where foo and the variable flag are used
(the block using testLoopCount and the subsequent flag = true; foo(...) call)
and wrap or replace the final invocation with an assertion that expects an
exception from foo(f, [10,20,30]) to ensure the flagged exception path is
exercised.

In `@test/js/bun/jsc-stress/fixtures/wasm/bbq-osr-with-exceptions.js`:
- Around line 29-35: The catch block around the loop that calls
fn2(global1.value, global11.value, fn3) currently swallows all exceptions;
change it so only RangeError is handled silently and any other exception is
rethrown (or otherwise propagated). Locate the try/catch that wraps the for-loop
calling fn2 and update the catch to check e instanceof RangeError and rethrow e
when it is not a RangeError.

In `@test/js/bun/jsc-stress/fixtures/wasm/ipint-bbq-osr-with-try4.js`:
- Around line 33-36: The test fixture uses a high-precision float literal in fn2
(8.433544074882645e1) that will lose precision at runtime; open the fn2 function
and replace the scientific notation literal with a runtime-stable value (e.g.,
84.33544074882646 or a shortened 84.33544) so the stored test data matches the
value produced at runtime and avoids misleading precision differences.

In `@test/js/bun/jsc-stress/fixtures/wasm/ipint-bbq-osr-with-try5.js`:
- Around line 2-9: Normalize inconsistent indentation around the helper setup:
adjust the indentation of the instantiate function body and the subsequent
declarations (instantiate, bytes, report, isJIT, extra, and the async IIFE) so
they match the file's prevailing style (make the function body and inner lines
use the same indentation as the surrounding top-level declarations) to keep
formatting consistent.
- Around line 1-12: Add the same skip-mode directives used in the other variants
to this test so it will be skipped when JIT is disabled: insert the two lines
that push "wasm-no-jit" and "wasm-no-wasm-jit" into $skipModes near the top of
the file (above or just before the instantiate function) so the test with
symbols like instantiate, callerIsBBQOrOMGCompiled, and extra={isJIT} will be
skipped in non-JIT runs.

In `@test/js/bun/jsc-stress/fixtures/wasm/omg-recompile-from-two-bbq.js`:
- Around line 27-34: The try/catch and promise handlers currently swallow errors
and provide no progress reporting; update the catch block that wraps the loop
invoking fn0(fn1) to call report('error', e) (or report('error') plus logging
the exception) instead of being empty, and add a .then(() => report('after'))
and a terminal .catch(e => report('error', e)) on the immediately-invoked async
call so both success and failure are reported; refer to the symbols fn0, fn1,
report and the IIFE promise tail in the diff when making the changes.
- Around line 25-26: The JSDoc type annotation for the destructured exports is
incomplete—replace the truncated comment around (i0.instance.exports) with a
full `@type` block so IDEs recognize types; update the comment that surrounds the
destructuring of fn0, fn1, global5..global8, table3..table13, tag0..tag3 to
start with /** `@type` {{ ... }} */ and list the exported symbols (fn0, fn1,
global5, global6, global7, global8, table3, table4, table5, table6, table7,
table8, table9, table10, table11, table12, table13, tag0, tag1, tag2, tag3} so
the annotation is well-formed and not truncated, leaving the destructuring
assignment (i0.instance.exports) unchanged.

In `@test/js/bun/json5/generate_json5_test_suite.ts`:
- Around line 54-56: The hardcoded relative module path in json5PkgPath is
fragile; update the loader for JSON5 so it first attempts to require('json5')
normally and if that fails uses a robust resolve via require.resolve with
explicit search paths (e.g., process.cwd() or a repo-root 'bench/json5' path) to
locate the package, then assign the resolved path to json5PkgPath and load JSON5
into JSON5Ref; change references in generate_json5_test_suite.ts (json5PkgPath,
JSON5Ref) to use this try/catch + require.resolve fallback instead of the fixed
"../../../../bench/json5/node_modules/json5" string.

In `@test/js/bun/json5/json5.test.ts`:
- Around line 1183-1186: In the "deeply nested arrays" test replace use of
String.prototype.repeat for building the input with Buffer.alloc to follow repo
guidelines: instead of "[".repeat(depth) + "1" + "]".repeat(depth) construct the
opening and closing repeated segments using Buffer.alloc(depth, "[").toString()
and Buffer.alloc(depth, "]").toString() (keep the variables depth, input and
expected as-is) so the test builds the same nested string without using
.repeat().

Comment thread bench/json5/json5.mjs
Comment on lines +12 to +57
// -- parse inputs --

const smallJson5 = `{
// User profile
name: "John Doe",
age: 30,
email: 'john@example.com',
active: true,
}`;

function generateLargeJson5(count) {
const lines = ["{\n // Auto-generated dataset\n items: [\n"];
for (let i = 0; i < count; i++) {
lines.push(` {
id: ${i},
name: 'item_${i}',
value: ${(Math.random() * 1000).toFixed(2)},
hex: 0x${i.toString(16).toUpperCase()},
active: ${i % 2 === 0},
tags: ['tag_${i % 10}', 'category_${i % 5}',],
// entry ${i}
},\n`);
}
lines.push(" ],\n total: " + count + ",\n status: 'complete',\n}\n");
return lines.join("");
}

const largeJson5 = generateLargeJson5(6500);

// -- stringify inputs --

const smallObject = {
name: "John Doe",
age: 30,
email: "john@example.com",
active: true,
};

const largeObject = {
items: Array.from({ length: 10000 }, (_, i) => ({
id: i,
name: `item_${i}`,
value: +(Math.random() * 1000).toFixed(2),
active: i % 2 === 0,
tags: [`tag_${i % 10}`, `category_${i % 5}`],
})),

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🧹 Nitpick | 🔵 Trivial

Make benchmark data deterministic for stable comparisons.

Random inputs add variance between runs; a fixed PRNG seed keeps results reproducible.

♻️ Proposed refactor
+function mulberry32(seed) {
+  return function () {
+    let t = (seed += 0x6d2b79f5);
+    t = Math.imul(t ^ (t >>> 15), t | 1);
+    t ^= t + Math.imul(t ^ (t >>> 7), t | 61);
+    return ((t ^ (t >>> 14)) >>> 0) / 4294967296;
+  };
+}
+const rand = mulberry32(0x1a2b3c4d);
+
 function generateLargeJson5(count) {
   const lines = ["{\n  // Auto-generated dataset\n  items: [\n"];
   for (let i = 0; i < count; i++) {
     lines.push(`    {
       id: ${i},
       name: 'item_${i}',
-      value: ${(Math.random() * 1000).toFixed(2)},
+      value: ${(rand() * 1000).toFixed(2)},
       hex: 0x${i.toString(16).toUpperCase()},
       active: ${i % 2 === 0},
       tags: ['tag_${i % 10}', 'category_${i % 5}',],
       // entry ${i}
     },\n`);
   }
   lines.push("  ],\n  total: " + count + ",\n  status: 'complete',\n}\n");
   return lines.join("");
 }
@@
 const largeObject = {
   items: Array.from({ length: 10000 }, (_, i) => ({
     id: i,
     name: `item_${i}`,
-    value: +(Math.random() * 1000).toFixed(2),
+    value: +(rand() * 1000).toFixed(2),
     active: i % 2 === 0,
     tags: [`tag_${i % 10}`, `category_${i % 5}`],
   })),
   total: 10000,
   status: "complete",
 };
🤖 Prompt for AI Agents
In `@bench/json5/json5.mjs` around lines 12 - 57, The benchmark data is
non-deterministic because Math.random() is used in generateLargeJson5 and when
building largeObject; replace those Math.random() calls with a seeded PRNG so
runs are reproducible: add a small deterministic generator (e.g.,
mulberry32/xorshift) initialized with a fixed seed and call it instead of
Math.random() inside generateLargeJson5 (for value and hex generation) and
inside the largeObject Array.from mapping (for value), ensuring hex formatting
still uses the PRNG output and all places that used Math.random() are switched
to the seeded generator (refer to function generateLargeJson5 and the
largeObject constant).

The Bun bundler implements a set of default loaders out of the box. As a rule of thumb, the bundler and the runtime both support the same set of file types out of the box.

`.js` `.cjs` `.mjs` `.mts` `.cts` `.ts` `.tsx` `.jsx` `.css` `.json` `.jsonc` `.toml` `.yaml` `.yml` `.txt` `.wasm` `.node` `.html` `.sh`
`.js` `.cjs` `.mjs` `.mts` `.cts` `.ts` `.tsx` `.jsx` `.css` `.json` `.jsonc` `.json5` `.toml` `.yaml` `.yml` `.txt` `.wasm` `.node` `.html` `.sh`

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

⚠️ Potential issue | 🟡 Minor

🧩 Analysis chain

🏁 Script executed:

#!/bin/bash
# First, let's check if the schema.zig file exists and search for Loader enum
find . -name "schema.zig" -type f

Repository: oven-sh/bun

Length of output: 102


🏁 Script executed:

#!/bin/bash
# Search for Loader enum and md loader references
rg -n "enum.*Loader|type.*Loader" --type zig
rg -n '\.md|"md"' --type zig | head -30

Repository: oven-sh/bun

Length of output: 3788


🏁 Script executed:

#!/bin/bash
# Check the documentation file mentioned in the review
cat -n docs/runtime/file-types.mdx

Repository: oven-sh/bun

Length of output: 16195


Add .md to the default extension list.

The .md loader is implemented and actively used throughout the codebase (transpiler, options parsing), but it's missing from the overview list. This omission on line 8 suggests the loader isn't supported, which is misleading.

Proposed doc fix
-`.js` `.cjs` `.mjs` `.mts` `.cts` `.ts` `.tsx` `.jsx` `.css` `.json` `.jsonc` `.json5` `.toml` `.yaml` `.yml` `.txt` `.wasm` `.node` `.html` `.sh`
+`.js` `.cjs` `.mjs` `.mts` `.cts` `.ts` `.tsx` `.jsx` `.css` `.json` `.jsonc` `.json5` `.md` `.toml` `.yaml` `.yml` `.txt` `.wasm` `.node` `.html` `.sh`
🤖 Prompt for AI Agents
In `@docs/runtime/file-types.mdx` at line 8, The docs list of supported file
extensions is missing `.md`, which is misleading because the `.md` loader
exists; update the extensions list in the file-types overview by adding `.md` to
the existing inline list (the string containing `.js .cjs .mjs ... .sh`) so the
markdown loader is documented alongside the other extensions.

Comment on lines +1149 to +1192
interface RenderCallbacks {
/** Heading (level 1–6). `id` is set when `headings: { ids: true }` is enabled. */
heading?: (children: string, meta: HeadingMeta) => string | null | undefined;
/** Paragraph. */
paragraph?: (children: string) => string | null | undefined;
/** Blockquote. */
blockquote?: (children: string) => string | null | undefined;
/** Code block. `meta.language` is the info-string (e.g. `"js"`). Only passed for fenced code blocks with a language. */
code?: (children: string, meta?: CodeBlockMeta) => string | null | undefined;
/** Ordered or unordered list. `start` is the first item number for ordered lists. */
list?: (children: string, meta: ListMeta) => string | null | undefined;
/** List item. `meta.checked` is set for task list items (`- [x]` / `- [ ]`). Only passed for task list items. */
listItem?: (children: string, meta?: ListItemMeta) => string | null | undefined;
/** Horizontal rule. */
hr?: (children: string) => string | null | undefined;
/** Table. */
table?: (children: string) => string | null | undefined;
/** Table head. */
thead?: (children: string) => string | null | undefined;
/** Table body. */
tbody?: (children: string) => string | null | undefined;
/** Table row. */
tr?: (children: string) => string | null | undefined;
/** Table header cell. `meta.align` is set when column alignment is specified. */
th?: (children: string, meta?: CellMeta) => string | null | undefined;
/** Table data cell. `meta.align` is set when column alignment is specified. */
td?: (children: string, meta?: CellMeta) => string | null | undefined;
/** Raw HTML content. */
html?: (children: string) => string | null | undefined;
/** Strong emphasis (`**text**`). */
strong?: (children: string) => string | null | undefined;
/** Emphasis (`*text*`). */
emphasis?: (children: string) => string | null | undefined;
/** Link. `href` is the URL, `title` is the optional title attribute. */
link?: (children: string, meta: LinkMeta) => string | null | undefined;
/** Image. `src` is the URL, `title` is the optional title attribute. */
image?: (children: string, meta: ImageMeta) => string | null | undefined;
/** Inline code (`` `code` ``). */
codespan?: (children: string) => string | null | undefined;
/** Strikethrough (`~~text~~`). */
strikethrough?: (children: string) => string | null | undefined;
/** Plain text content. */
text?: (text: string) => string | null | undefined;
}

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

⚠️ Potential issue | 🟡 Minor

🧩 Analysis chain

🏁 Script executed:

# Find markdown-related files in the repository
fd -i markdown -type f | head -20

Repository: oven-sh/bun

Length of output: 225


🏁 Script executed:

# List files in the packages/bun-types directory to understand the structure
git ls-files packages/bun-types/ | head -20

Repository: oven-sh/bun

Length of output: 692


🏁 Script executed:

# Search for RenderCallbacks definition to see current state
rg -n "interface RenderCallbacks" packages/bun-types/

Repository: oven-sh/bun

Length of output: 119


🏁 Script executed:

# Search for math and underline references in bun.d.ts
rg -n "math|underline" packages/bun-types/bun.d.ts

Repository: oven-sh/bun

Length of output: 326


🏁 Script executed:

# Read the RenderCallbacks interface section
sed -n '1149,1192p' packages/bun-types/bun.d.ts

Repository: oven-sh/bun

Length of output: 2724


🏁 Script executed:

# Check the React overrides section to see if math/u are there
sed -n '1070,1110p' packages/bun-types/bun.d.ts

Repository: oven-sh/bun

Length of output: 1509


🏁 Script executed:

# Search for where math callback might be defined or used in render context
rg -n "math" packages/bun-types/bun.d.ts -B2 -A2

Repository: oven-sh/bun

Length of output: 591


🏁 Script executed:

# Search for underline callback
rg -n "underline" packages/bun-types/bun.d.ts -B2 -A2

Repository: oven-sh/bun

Length of output: 441


🏁 Script executed:

# Find markdown-related implementation files
find . -name "*markdown*" -o -name "*render*" | grep -v node_modules | grep -v ".git" | head -20

Repository: oven-sh/bun

Length of output: 351


🏁 Script executed:

# Search for Zig files that might contain the markdown implementation
fd "\.zig$" | grep -i markdown

Repository: oven-sh/bun

Length of output: 88


🏁 Script executed:

# Broader search for markdown implementation
rg -l "math|underline" --type zig | head -10

Repository: oven-sh/bun

Length of output: 247


🏁 Script executed:

# Search for where render() callback is called with these nodes
rg -n "render.*math|render.*underline" | head -20

Repository: oven-sh/bun

Length of output: 50368


🏁 Script executed:

# Search for math and underline handling in render_blocks.zig
rg -n "math|underline" src/md/render_blocks.zig -A3 -B3

Repository: oven-sh/bun

Length of output: 425


🏁 Script executed:

# Search in html_renderer.zig
rg -n "math|underline" src/md/html_renderer.zig -A3 -B3

Repository: oven-sh/bun

Length of output: 834


🏁 Script executed:

# Check the test file for render callbacks
cat test/js/bun/md/md-render-callback.test.ts

Repository: oven-sh/bun

Length of output: 9779


🏁 Script executed:

# Search for how callbacks are invoked in the render implementation
rg -n "RenderCallbacks" packages/bun-types/bun.d.ts -B5 -A25 | head -80

Repository: oven-sh/bun

Length of output: 3421


🏁 Script executed:

# Look for where math/underline nodes might be emitted in zig
rg -n "\.latexmath|\.u\b" src/md/ -B2 -A2

Repository: oven-sh/bun

Length of output: 1473


🏁 Script executed:

# Check MarkdownObject.zig for callback handling
rg -n "math|underline" src/bun.js/api/MarkdownObject.zig -B3 -A3

Repository: oven-sh/bun

Length of output: 1437


🏁 Script executed:

# Search for the callback renderer implementation
fd "callback" src/md/ -type f

Repository: oven-sh/bun

Length of output: 225


🏁 Script executed:

# Look for where render() calls callbacks for different node types
rg -n "heading|paragraph|strong" src/md/ -l

Repository: oven-sh/bun

Length of output: 270


🏁 Script executed:

# Search in MarkdownObject for the render() callback implementation
rg -n "render.*callback\|RenderCallback" src/bun.js/api/MarkdownObject.zig -B5 -A15 | head -100

Repository: oven-sh/bun

Length of output: 37


🏁 Script executed:

# Look for where callbacks are invoked - search for callback function calls
rg -n "callback" src/bun.js/api/MarkdownObject.zig | head -30

Repository: oven-sh/bun

Length of output: 2056


🏁 Script executed:

# Search for how render() distinguishes between callback and component rendering
rg -n "render\|callback" src/bun.js/api/MarkdownObject.zig -A2 -B2 | head -80

Repository: oven-sh/bun

Length of output: 37


🏁 Script executed:

# Check if there's a separate callback renderer
rg -n "struct.*Callback\|CallbackRenderer" src/bun.js/api/ -A10

Repository: oven-sh/bun

Length of output: 37


🏁 Script executed:

# Look at MarkdownObject.zig structure to understand the render flow
head -100 src/bun.js/api/MarkdownObject.zig

Repository: oven-sh/bun

Length of output: 4320


🏁 Script executed:

# Look at the Callbacks struct and how callbacks are extracted
sed -n '746,900p' src/bun.js/api/MarkdownObject.zig

Repository: oven-sh/bun

Length of output: 6037


🏁 Script executed:

# Look for getBlockCallback and getSpanCallback to see what callbacks are supported
rg -n "getBlockCallback\|getSpanCallback" src/bun.js/api/MarkdownObject.zig -A20

Repository: oven-sh/bun

Length of output: 37


🏁 Script executed:

# Check the Callbacks struct definition
rg -n "struct.*Callbacks\|#callbacks" src/bun.js/api/MarkdownObject.zig -B2 -A30 | head -80

Repository: oven-sh/bun

Length of output: 37


🏁 Script executed:

# Search for getBlockCallback function
rg -n "fn getBlockCallback" src/bun.js/api/MarkdownObject.zig -A30

Repository: oven-sh/bun

Length of output: 1410


🏁 Script executed:

# Search for getSpanCallback function  
rg -n "fn getSpanCallback" src/bun.js/api/MarkdownObject.zig -A30

Repository: oven-sh/bun

Length of output: 1558


🏁 Script executed:

# Check what all the callback functions look like
sed -n '900,980p' src/bun.js/api/MarkdownObject.zig

Repository: oven-sh/bun

Length of output: 3638


🏁 Script executed:

# Confirm that .u and .latexmath are not handled in getSpanCallback
sed -n '1009,1020p' src/bun.js/api/MarkdownObject.zig

Repository: oven-sh/bun

Length of output: 492


Add missing math and underline callbacks to RenderCallbacks interface.

The latexMath and underline options enable parsing of these markdown features, and React overrides (math, u components) handle them. However, RenderCallbacks lacks corresponding callbacks. When render() is called with these options enabled, the parsed math/underline nodes are silently skipped because getSpanCallback() returns .zero for unmapped span types. Either add the callbacks so users can customize rendering, or document that these features work only with html() and react(), not render().

Suggested typing addition
       /** Strikethrough (`~~text~~`). */
       strikethrough?: (children: string) => string | null | undefined;
+      /** LaTeX math (`$...$` / `$$...$$`). */
+      math?: (children: string) => string | null | undefined;
+      /** Underline (`__text__`). */
+      underline?: (children: string) => string | null | undefined;
       /** Plain text content. */
       text?: (text: string) => string | null | undefined;
🤖 Prompt for AI Agents
In `@packages/bun-types/bun.d.ts` around lines 1149 - 1192, RenderCallbacks is
missing handlers for parsed math and underline spans, so when
latexMath/underline are enabled render() silently drops those nodes; add two
callbacks to the RenderCallbacks interface—math (e.g. math?: (children: string,
meta?: MathMeta) => string | null | undefined) and underline (e.g. underline?:
(children: string) => string | null | undefined)—so getSpanCallback()/render()
can map those span types instead of returning .zero; update types referenced
(MathMeta) or reuse existing meta shapes as appropriate and ensure
RenderCallbacks includes these new symbols.

Comment thread src/api/schema.zig
Comment on lines 345 to 348
yaml = 19,
json5 = 20,
md = 21,
_,

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

⚠️ Potential issue | 🔴 Critical

🧩 Analysis chain

🏁 Script executed:

#!/bin/bash
# Verify loader enum values and C/C++ mapping consistency
rg -n "BunLoaderType|Loader.*enum|json5|md" -g '!**/third_party/**' | head -100

Repository: oven-sh/bun

Length of output: 11038


🏁 Script executed:

#!/bin/bash
# Search for Loader enum in schema.zig
echo "=== Loader enum in schema.zig ==="
rg -A 30 "pub const Loader.*enum" src/api/schema.zig

echo -e "\n=== C/C++ BunLoaderType ==="
find . -name "headers-handwritten.h" -o -name "*loader*.h" | head -20

echo -e "\n=== Search for BunLoaderType definition ==="
rg "BunLoaderType|enum.*Loader" -g "*.h" -g "*.hpp" 2>/dev/null | head -50

Repository: oven-sh/bun

Length of output: 2404


🏁 Script executed:

#!/bin/bash
# Get all BunLoaderType definitions from headers
echo "=== All BunLoaderType constants in headers-handwritten.h ==="
rg "const BunLoaderType.*=" src/bun.js/bindings/headers-handwritten.h -o | sort

echo -e "\n=== Check for json5 definition ==="
rg "json5|JSON5" src/bun.js/bindings/headers-handwritten.h

echo -e "\n=== Full loader section from headers ==="
rg -B 2 -A 25 "BunLoaderTypeNone" src/bun.js/bindings/headers-handwritten.h | head -60

Repository: oven-sh/bun

Length of output: 1783


🏁 Script executed:

#!/bin/bash
# Check for loader string mappings
echo "=== Search for loader string mappings ==="
rg -i "loaderToString|loader.*string|getLoaderName" -g "*.h" -g "*.cpp" -g "*.ts" -g "*.zig" 2>/dev/null | head -40

echo -e "\n=== Check what's between NAPI(11) and YAML(19) in C headers ==="
rg "BunLoaderType.*=" src/bun.js/bindings/headers-handwritten.h | grep -E "(12|13|14|15|16|17|18)"

echo -e "\n=== Check full zig enum for context on base64, dataurl, text, bunsh, sqlite ==="
sed -n '340,365p' src/api/schema.zig

Repository: oven-sh/bun

Length of output: 50367


🏁 Script executed:

#!/bin/bash
# Get complete loader section from headers
sed -n '/Must be kept in sync with bun.schema.api.Loader/,/pragma mark - Stream/p' src/bun.js/bindings/headers-handwritten.h

Repository: oven-sh/bun

Length of output: 773


Add missing loader type constants to C/C++ bindings and resolve numeric conflicts.

The Zig Loader enum and C/C++ BunLoaderType definitions are out of sync despite the explicit requirement in headers-handwritten.h. The C/C++ headers are missing constants for values 12–18 (base64, dataurl, text, bunsh, sqlite, sqlite_embedded, html), json5 is completely absent, and md has conflicting values (21 in Zig vs. 20 in C/C++). Add all missing BunLoaderType constants and ensure numeric values align precisely with the Zig enum.

🤖 Prompt for AI Agents
In `@src/api/schema.zig` around lines 345 - 348, The C/C++ binding enum
BunLoaderType in headers-handwritten.h is missing several members and has
numeric mismatches with the Zig Loader enum in src/api/schema.zig; add constants
for base64, dataurl, text, bunsh, sqlite, sqlite_embedded, html, and json5 with
the exact integer values that match the Zig Loader sequence (fill in values
12–18 for the missing ones and add json5 at its correct value), and update the
md constant to use the same numeric value as the Zig enum (ensure all
BunLoaderType numeric assignments mirror the Zig Loader order so there are no
conflicts between md and json5).

Comment on lines +2106 to 2109
auto arrayBuffer = ArrayBuffer::createFromBytes({ reinterpret_cast<const uint8_t*>(external_data), byte_length }, createSharedTask<void(void*)>([env = WTF::Ref<NapiEnv>(*env), finalize_hint, finalize_cb](void* p) {
NAPI_LOG("external ArrayBuffer finalizer");
env->doFinalizer(finalize_cb, p, finalize_hint);
}));

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

⚠️ Potential issue | 🟠 Major

🧩 Analysis chain

🏁 Script executed:

# First, let's locate the JSArrayBuffer::create implementation
rg -n "JSArrayBuffer::create" --type cpp --type h | head -20

Repository: oven-sh/bun

Length of output: 3867


🏁 Script executed:

# Also search for the specific file and lines mentioned
wc -l src/bun.js/bindings/napi.cpp

Repository: oven-sh/bun

Length of output: 88


🏁 Script executed:

# Read the context around lines 2106-2109 in napi.cpp
sed -n '2090,2130p' src/bun.js/bindings/napi.cpp | cat -n

Repository: oven-sh/bun

Length of output: 1807


🏁 Script executed:

# Find JSArrayBuffer::create definition
rg "class JSArrayBuffer" --type cpp --type h -A 20 | head -50

Repository: oven-sh/bun

Length of output: 1069


🏁 Script executed:

# Look for exception handling in napi.cpp - check NAPI_PREAMBLE and similar
grep -n "NAPI_PREAMBLE\|ThrowScope\|ExceptionScope" src/bun.js/bindings/napi.cpp | head -20

Repository: oven-sh/bun

Length of output: 943


🏁 Script executed:

# Get broader context around the function, including earlier lines
sed -n '2090,2150p' src/bun.js/bindings/napi.cpp | cat -n

Repository: oven-sh/bun

Length of output: 2651


🏁 Script executed:

# Find the full NAPI_PREAMBLE macro definition
sed -n '87,130p' src/bun.js/bindings/napi.cpp | cat -n

Repository: oven-sh/bun

Length of output: 2729


🏁 Script executed:

# Search for exception handling macros and how they work in napi.cpp
grep -n "NAPI_RETURN_IF_EXCEPTION\|NAPI_CHECK_EXCEPTION" src/bun.js/bindings/napi.cpp | head -20

Repository: oven-sh/bun

Length of output: 1021


🏁 Script executed:

# Look at how JSArrayBuffer::create is used in other napi.cpp functions
rg -B5 -A5 "JSArrayBuffer::create" src/bun.js/bindings/napi.cpp

Repository: oven-sh/bun

Length of output: 1063


Add exception handling after JSArrayBuffer::create to prevent finalizer execution on creation failure.

JSArrayBuffer::create can throw exceptions (as shown by the exception check at line 649 in the same file), but line 2111 lacks the NAPI_RETURN_IF_EXCEPTION(env) check present in that similar code. If creation fails, the createSharedTask finalizer will still execute and call finalize_cb on caller-owned memory, causing unexpected deallocation.

Add NAPI_RETURN_IF_EXCEPTION(env); immediately after the JSArrayBuffer::create call:

Fix
     auto arrayBuffer = ArrayBuffer::createFromBytes({ reinterpret_cast<const uint8_t*>(external_data), byte_length }, createSharedTask<void(void*)>([env = WTF::Ref<NapiEnv>(*env), finalize_hint, finalize_cb](void* p) {
         NAPI_LOG("external ArrayBuffer finalizer");
         env->doFinalizer(finalize_cb, p, finalize_hint);
     }));
 
     auto* buffer = JSC::JSArrayBuffer::create(vm, globalObject->arrayBufferStructure(ArrayBufferSharingMode::Default), WTF::move(arrayBuffer));
+    NAPI_RETURN_IF_EXCEPTION(env);
 
     *result = toNapi(buffer, globalObject);
     NAPI_RETURN_SUCCESS(env);
🤖 Prompt for AI Agents
In `@src/bun.js/bindings/napi.cpp` around lines 2106 - 2109, The ArrayBuffer
creation using ArrayBuffer::createFromBytes with createSharedTask can throw, but
the code does not guard with NAPI_RETURN_IF_EXCEPTION(env), so if creation fails
the shared finalizer may run and call finalize_cb on caller-owned memory;
immediately after the ArrayBuffer::createFromBytes(...) call (the line
constructing arrayBuffer) add NAPI_RETURN_IF_EXCEPTION(env) to return on
exception and prevent scheduling the createSharedTask finalizer; reference
symbols: ArrayBuffer::createFromBytes, createSharedTask, finalize_cb,
finalize_hint, env.

Comment on lines +2 to +9
function instantiate(moduleBase64, importObject) {
let bytes = Uint8Array.fromBase64(moduleBase64);
return WebAssembly.instantiate(bytes, importObject);
}
const report = $.agent.report;
const isJIT = callerIsBBQOrOMGCompiled;
const extra = {isJIT};
(async function () {

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🧹 Nitpick | 🔵 Trivial

Minor: Inconsistent indentation in helper setup.

The instantiate function body and early declarations have mixed indentation (some lines with 4 spaces, others without). This is cosmetic but inconsistent with the rest of the file.

🤖 Prompt for AI Agents
In `@test/js/bun/jsc-stress/fixtures/wasm/ipint-bbq-osr-with-try5.js` around lines
2 - 9, Normalize inconsistent indentation around the helper setup: adjust the
indentation of the instantiate function body and the subsequent declarations
(instantiate, bytes, report, isJIT, extra, and the async IIFE) so they match the
file's prevailing style (make the function body and inner lines use the same
indentation as the surrounding top-level declarations) to keep formatting
consistent.

Comment on lines +25 to +26
let {fn0, fn1, global5, global6, global7, global8, table3, table4, table5, table6, table7, table8, table9, table10, table11, table12, table13, tag0, tag1, tag2, tag3} = /**
}} */ (i0.instance.exports);

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

⚠️ Potential issue | 🟡 Minor

Incomplete JSDoc type annotation.

The JSDoc type definition is truncated—it has the closing }} */ but missing the opening @type {{. This may cause IDE warnings.

📝 Proposed fix
 let {fn0, fn1, global5, global6, global7, global8, table3, table4, table5, table6, table7, table8, table9, table10, table11, table12, table13, tag0, tag1, tag2, tag3} = /**
-  }} */ (i0.instance.exports);
+  `@type` {{
+fn0: (a0: FuncRef) => I32,
+fn1: () => void,
+global5: WebAssembly.Global,
+global6: WebAssembly.Global,
+global7: WebAssembly.Global,
+global8: WebAssembly.Global,
+table3: WebAssembly.Table,
+table4: WebAssembly.Table,
+table5: WebAssembly.Table,
+table6: WebAssembly.Table,
+table7: WebAssembly.Table,
+table8: WebAssembly.Table,
+table9: WebAssembly.Table,
+table10: WebAssembly.Table,
+table11: WebAssembly.Table,
+table12: WebAssembly.Table,
+table13: WebAssembly.Table,
+tag0: WebAssembly.Tag,
+tag1: WebAssembly.Tag,
+tag2: WebAssembly.Tag,
+tag3: WebAssembly.Tag
+  }} */ (i0.instance.exports);
📝 Committable suggestion

‼️ IMPORTANT
Carefully review the code before committing. Ensure that it accurately replaces the highlighted code, contains no missing lines, and has no issues with indentation. Thoroughly test & benchmark the code to ensure it meets the requirements.

Suggested change
let {fn0, fn1, global5, global6, global7, global8, table3, table4, table5, table6, table7, table8, table9, table10, table11, table12, table13, tag0, tag1, tag2, tag3} = /**
}} */ (i0.instance.exports);
let {fn0, fn1, global5, global6, global7, global8, table3, table4, table5, table6, table7, table8, table9, table10, table11, table12, table13, tag0, tag1, tag2, tag3} = /**
`@type` {{
fn0: (a0: FuncRef) => I32,
fn1: () => void,
global5: WebAssembly.Global,
global6: WebAssembly.Global,
global7: WebAssembly.Global,
global8: WebAssembly.Global,
table3: WebAssembly.Table,
table4: WebAssembly.Table,
table5: WebAssembly.Table,
table6: WebAssembly.Table,
table7: WebAssembly.Table,
table8: WebAssembly.Table,
table9: WebAssembly.Table,
table10: WebAssembly.Table,
table11: WebAssembly.Table,
table12: WebAssembly.Table,
table13: WebAssembly.Table,
tag0: WebAssembly.Tag,
tag1: WebAssembly.Tag,
tag2: WebAssembly.Tag,
tag3: WebAssembly.Tag
}} */ (i0.instance.exports);
🤖 Prompt for AI Agents
In `@test/js/bun/jsc-stress/fixtures/wasm/omg-recompile-from-two-bbq.js` around
lines 25 - 26, The JSDoc type annotation for the destructured exports is
incomplete—replace the truncated comment around (i0.instance.exports) with a
full `@type` block so IDEs recognize types; update the comment that surrounds the
destructuring of fn0, fn1, global5..global8, table3..table13, tag0..tag3 to
start with /** `@type` {{ ... }} */ and list the exported symbols (fn0, fn1,
global5, global6, global7, global8, table3, table4, table5, table6, table7,
table8, table9, table10, table11, table12, table13, tag0, tag1, tag2, tag3} so
the annotation is well-formed and not truncated, leaving the destructuring
assignment (i0.instance.exports) unchanged.

Comment on lines +27 to +34
try {
for (let k=0; k<24; k++) {
let zzz = fn0(fn1);
}
} catch (e) {
}
})().then(() => {
}) No newline at end of file

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🧹 Nitpick | 🔵 Trivial

Missing error handling and progress reporting.

Unlike other wasm stress fixtures in this PR (e.g., ipint-bbq-osr-check-try-implicit-slot-overlap.js), this fixture:

  1. Has an empty catch block that silently swallows all exceptions
  2. Has empty .then() and no .catch() handlers—no report('after') or report('error') calls

This inconsistency may make debugging failures harder and reduces observability.

📝 Proposed fix
+report('progress');
 try {
   for (let k=0; k<24; k++) {
   let zzz = fn0(fn1);
   }
 } catch (e) {
+  if (e instanceof WebAssembly.Exception) {
+  } else if (e instanceof TypeError) {
+    if (e.message === 'an exported wasm function cannot contain a v128 parameter or return value') {} else { throw e; }
+  } else if (e instanceof WebAssembly.RuntimeError || e instanceof RangeError) {} else { throw e; }
 }
 })().then(() => {
+  report('after');
+}).catch(e => {
+  report('error');
 })
🤖 Prompt for AI Agents
In `@test/js/bun/jsc-stress/fixtures/wasm/omg-recompile-from-two-bbq.js` around
lines 27 - 34, The try/catch and promise handlers currently swallow errors and
provide no progress reporting; update the catch block that wraps the loop
invoking fn0(fn1) to call report('error', e) (or report('error') plus logging
the exception) instead of being empty, and add a .then(() => report('after'))
and a terminal .catch(e => report('error', e)) on the immediately-invoked async
call so both success and failure are reported; refer to the symbols fn0, fn1,
report and the IIFE promise tail in the diff when making the changes.

Comment on lines +54 to +56
// The json5 npm package – resolve from bench/json5 where it's installed
const json5PkgPath = join(import.meta.dir, "../../../../bench/json5/node_modules/json5");
const JSON5Ref = require(json5PkgPath) as { parse: (s: string) => unknown };

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🧹 Nitpick | 🔵 Trivial

Relative path to bench/json5 may be fragile.

The hardcoded relative path ../../../../bench/json5/node_modules/json5 depends on the file's location in the directory tree. If the generator script is moved, this path will break.

Consider using a more robust approach:

♻️ Suggested improvement
-const json5PkgPath = join(import.meta.dir, "../../../../bench/json5/node_modules/json5");
-const JSON5Ref = require(json5PkgPath) as { parse: (s: string) => unknown };
+// Find repo root and resolve json5 from there
+import { dirname } from "node:path";
+function findRepoRoot(start: string): string {
+  let dir = start;
+  while (dir !== dirname(dir)) {
+    if (existsSync(join(dir, "package.json")) && existsSync(join(dir, "bench"))) {
+      return dir;
+    }
+    dir = dirname(dir);
+  }
+  throw new Error("Could not find repository root");
+}
+const repoRoot = findRepoRoot(import.meta.dir);
+const json5PkgPath = join(repoRoot, "bench/json5/node_modules/json5");
+const JSON5Ref = require(json5PkgPath) as { parse: (s: string) => unknown };
🤖 Prompt for AI Agents
In `@test/js/bun/json5/generate_json5_test_suite.ts` around lines 54 - 56, The
hardcoded relative module path in json5PkgPath is fragile; update the loader for
JSON5 so it first attempts to require('json5') normally and if that fails uses a
robust resolve via require.resolve with explicit search paths (e.g.,
process.cwd() or a repo-root 'bench/json5' path) to locate the package, then
assign the resolved path to json5PkgPath and load JSON5 into JSON5Ref; change
references in generate_json5_test_suite.ts (json5PkgPath, JSON5Ref) to use this
try/catch + require.resolve fallback instead of the fixed
"../../../../bench/json5/node_modules/json5" string.

Comment on lines +1183 to +1186
test("deeply nested arrays", () => {
const depth = 100;
const input = "[".repeat(depth) + "1" + "]".repeat(depth);
let expected: any = 1;

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🛠️ Refactor suggestion | 🟠 Major

Prefer Buffer.alloc(...).toString() over .repeat() in tests.
The repo guideline discourages .repeat() for repetitive strings in tests.

♻️ Suggested change
-    const input = "[".repeat(depth) + "1" + "]".repeat(depth);
+    const input =
+      Buffer.alloc(depth, "[").toString() + "1" + Buffer.alloc(depth, "]").toString();
As per coding guidelines, use `Buffer.alloc(count, fill).toString()` for repetitive test strings.
📝 Committable suggestion

‼️ IMPORTANT
Carefully review the code before committing. Ensure that it accurately replaces the highlighted code, contains no missing lines, and has no issues with indentation. Thoroughly test & benchmark the code to ensure it meets the requirements.

Suggested change
test("deeply nested arrays", () => {
const depth = 100;
const input = "[".repeat(depth) + "1" + "]".repeat(depth);
let expected: any = 1;
test("deeply nested arrays", () => {
const depth = 100;
const input =
Buffer.alloc(depth, "[").toString() + "1" + Buffer.alloc(depth, "]").toString();
let expected: any = 1;
🤖 Prompt for AI Agents
In `@test/js/bun/json5/json5.test.ts` around lines 1183 - 1186, In the "deeply
nested arrays" test replace use of String.prototype.repeat for building the
input with Buffer.alloc to follow repo guidelines: instead of "[".repeat(depth)
+ "1" + "]".repeat(depth) construct the opening and closing repeated segments
using Buffer.alloc(depth, "[").toString() and Buffer.alloc(depth,
"]").toString() (keep the variables depth, input and expected as-is) so the test
builds the same nested string without using .repeat().

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 1

🤖 Fix all issues with AI agents
In `@docs/runtime/markdown.mdx`:
- Around line 27-60: The docs show inconsistent defaults for GFM features: the
paragraph and inline comments for Bun.unstable_markdown.html claim features
(tables, strikethrough, tasklists, autolinks) are enabled by default while the
options table lists them as `false`; pick the correct default and make all three
places consistent by either (A) setting the table defaults for `tables`,
`strikethrough`, `tasklists`, and `autolinks` to `true` to match the
paragraph/inline comments, or (B) changing the opening paragraph and inline
comments to say they are disabled by default if `false` is the intended state;
ensure you update the option keys `tables`, `strikethrough`, `tasklists`, and
`autolinks` in the table and the example call to Bun.unstable_markdown.html
accordingly.

Comment thread docs/runtime/markdown.mdx Outdated
Comment thread src/md/blocks.zig
Comment on lines +4 to +15
var line_buf: [2]Line = .{ .{}, .{} };
var line_idx: u1 = 0;
var off: OFF = 0;

try self.enterBlock(.doc, 0, 0);

while (off < self.size) {
const line = &line_buf[line_idx];

try self.analyzeLine(off, &off, &pivot_line, line);
try self.processLine(&pivot_line, line, &line_buf, &line_idx);
}

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

this line_buf seems to be unused. only one line is used at a time

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 1

🤖 Fix all issues with AI agents
In `@test/js/bun/md/md-spec.test.ts`:
- Around line 15-20: In parseSpecFile replace the direct string repetition used
for the fence variable (currently const fence = "`".repeat(32)) with a
Buffer-based allocation per guidelines; set fence using Buffer.alloc(32,
"`").toString() so the test uses Buffer.alloc(count, fill).toString() instead of
.repeat().

Comment on lines +15 to +20
function parseSpecFile(path: string): SpecExample[] {
const content = readFileSync(path, "utf8").replace(/\r\n?/g, "\n");
const lines = content.split("\n");
const examples: SpecExample[] = [];
const fence = "`".repeat(32);
let i = 0;

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🧹 Nitpick | 🔵 Trivial

Use Buffer.alloc for the repeated fence string.
This aligns with the test string guideline and avoids direct string repetition.

♻️ Proposed update
-  const fence = "`".repeat(32);
+  const fence = Buffer.alloc(32, "`").toString();
As per coding guidelines: Use `Buffer.alloc(count, fill).toString()` instead of `'A'.repeat(count)` to create repetitive strings in tests.
📝 Committable suggestion

‼️ IMPORTANT
Carefully review the code before committing. Ensure that it accurately replaces the highlighted code, contains no missing lines, and has no issues with indentation. Thoroughly test & benchmark the code to ensure it meets the requirements.

Suggested change
function parseSpecFile(path: string): SpecExample[] {
const content = readFileSync(path, "utf8").replace(/\r\n?/g, "\n");
const lines = content.split("\n");
const examples: SpecExample[] = [];
const fence = "`".repeat(32);
let i = 0;
function parseSpecFile(path: string): SpecExample[] {
const content = readFileSync(path, "utf8").replace(/\r\n?/g, "\n");
const lines = content.split("\n");
const examples: SpecExample[] = [];
const fence = Buffer.alloc(32, "`").toString();
let i = 0;
🤖 Prompt for AI Agents
In `@test/js/bun/md/md-spec.test.ts` around lines 15 - 20, In parseSpecFile
replace the direct string repetition used for the fence variable (currently
const fence = "`".repeat(32)) with a Buffer-based allocation per guidelines; set
fence using Buffer.alloc(32, "`").toString() so the test uses
Buffer.alloc(count, fill).toString() instead of .repeat().

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 1

🤖 Fix all issues with AI agents
In `@src/bun.js/api/MarkdownObject.zig`:
- Around line 627-640: The loop that builds alt text uses a silent catch on
appendSlice which swallows OOMs and can produce truncated alt text; in
function/loop referencing alt_buf, entry.children.getIndex, child.toSlice,
alt_buf.appendSlice, and props.put (then bun.String.createUTF8ForJS), replace
the silent catch {} with explicit error handling: either propagate the
allocation error with try so the build fails, or handle the failure explicitly
(e.g., log the allocation error and break/abort building alt_buf) and add a
comment documenting graceful degradation; ensure the chosen approach
consistently handles errors before calling props.put.

Comment on lines +627 to +640
var alt_buf = std.ArrayListUnmanaged(u8){};
defer alt_buf.deinit(bun.default_allocator);
for (0..len) |i| {
const child = try entry.children.getIndex(g, @truncate(i));
if (child.isString()) {
const str = try child.toSlice(g, bun.default_allocator);
defer str.deinit();
alt_buf.appendSlice(bun.default_allocator, str.slice()) catch {};
}
}
if (alt_buf.items.len > 0) {
props.put(g, ZigString.static("alt"), try bun.String.createUTF8ForJS(g, alt_buf.items));
}
}

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

⚠️ Potential issue | 🟡 Minor

🧩 Analysis chain

🏁 Script executed:

# First, check if commit 1c58726 exists and what it changed
git log --oneline | grep -i "1c58726" || echo "Searching for commit hash in log..."
git show 1c58726 -- src/bun.js/api/MarkdownObject.zig 2>/dev/null | head -100 || echo "Commit info not available"

Repository: oven-sh/bun

Length of output: 4386


🏁 Script executed:

# Check the current state of MarkdownObject.zig around the alt text handling
wc -l src/bun.js/api/MarkdownObject.zig

Repository: oven-sh/bun

Length of output: 93


🏁 Script executed:

# Read the actual file around lines 627-640 to see current state
sed -n '620,645p' src/bun.js/api/MarkdownObject.zig

Repository: oven-sh/bun

Length of output: 1247


🏁 Script executed:

# Search for all occurrences of "catch {}" in the MarkdownObject.zig to see patterns
rg -n "catch \{\}" src/bun.js/api/MarkdownObject.zig

Repository: oven-sh/bun

Length of output: 148


🏁 Script executed:

# Also check if there are other similar patterns in the file
rg -n "appendSlice.*catch" src/bun.js/api/MarkdownObject.zig

Repository: oven-sh/bun

Length of output: 148


🏁 Script executed:

# Look at more context around the alt text handling
sed -n '620,650p' src/bun.js/api/MarkdownObject.zig

Repository: oven-sh/bun

Length of output: 1426


🏁 Script executed:

# Check if there are other similar error handling patterns in this function
sed -n '550,750p' src/bun.js/api/MarkdownObject.zig | grep -n -B2 -A2 "catch"

Repository: oven-sh/bun

Length of output: 324


🏁 Script executed:

# Search for comments or patterns related to alt text handling
rg -B5 -A5 "alt.*buf|alt.*text" src/bun.js/api/MarkdownObject.zig | head -40

Repository: oven-sh/bun

Length of output: 1096


🏁 Script executed:

# Check what the appendSlice error type is
rg -B3 -A3 "fn appendSlice" --type zig

Repository: oven-sh/bun

Length of output: 7625


🏁 Script executed:

# Double-check by looking at the actual full diff of commit 1c587264
git show 1c587264 --stat

Repository: oven-sh/bun

Length of output: 1604


🏁 Script executed:

# Check if there are any other commits that might have touched this area
git log --oneline --all -- src/bun.js/api/MarkdownObject.zig | head -20

Repository: oven-sh/bun

Length of output: 73


OOM still silently swallowed when building image alt text at line 634.

The code uses catch {} which silently drops allocation failures during alt text concatenation. If the allocator runs out of memory, the loop will continue with potentially incomplete alt_buf content.

The original claim that this was addressed in a previous commit is incorrect—commit 1c58726 refactored heading tracking but did not modify this code section.

Consider whether incomplete alt text should fail the build (use try) or be explicitly handled with a comment explaining the graceful degradation. As written now, it silently truncates accessibility content without indication.

🤖 Prompt for AI Agents
In `@src/bun.js/api/MarkdownObject.zig` around lines 627 - 640, The loop that
builds alt text uses a silent catch on appendSlice which swallows OOMs and can
produce truncated alt text; in function/loop referencing alt_buf,
entry.children.getIndex, child.toSlice, alt_buf.appendSlice, and props.put (then
bun.String.createUTF8ForJS), replace the silent catch {} with explicit error
handling: either propagate the allocation error with try so the build fails, or
handle the failure explicitly (e.g., log the allocation error and break/abort
building alt_buf) and add a comment documenting graceful degradation; ensure the
chosen approach consistently handles errors before calling props.put.

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 1

🤖 Fix all issues with AI agents
In `@src/md/render_blocks.zig`:
- Around line 74-83: The comment above the inner scan loop in render_row (the
while loop using start, end, row_text, and cell_index) incorrectly states it
skips "escaped chars and code spans" even though the loop only skips escaped
characters; update the comment to reflect current behavior (e.g., "Find cell
end, skipping escaped chars") and remove any mention of backtick/code-span
handling so it matches the implemented logic in that loop.

Comment thread src/md/render_blocks.zig
Comment on lines +74 to +83
while (start < row_text.len and cell_index < col_count) {
// Find cell end, skipping escaped chars and code spans
var end = start;
while (end < row_text.len and row_text[end] != '|') {
if (row_text[end] == '\\' and end + 1 < row_text.len) {
end += 2;
} else {
end += 1;
}
}

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

⚠️ Potential issue | 🟡 Minor

Update the comment about code-span handling.

The scan loop no longer skips backtick spans, so the comment is now inaccurate.

✏️ Suggested fix
-        // Find cell end, skipping escaped chars and code spans
+        // Find cell end, skipping escaped chars
📝 Committable suggestion

‼️ IMPORTANT
Carefully review the code before committing. Ensure that it accurately replaces the highlighted code, contains no missing lines, and has no issues with indentation. Thoroughly test & benchmark the code to ensure it meets the requirements.

Suggested change
while (start < row_text.len and cell_index < col_count) {
// Find cell end, skipping escaped chars and code spans
var end = start;
while (end < row_text.len and row_text[end] != '|') {
if (row_text[end] == '\\' and end + 1 < row_text.len) {
end += 2;
} else {
end += 1;
}
}
while (start < row_text.len and cell_index < col_count) {
// Find cell end, skipping escaped chars
var end = start;
while (end < row_text.len and row_text[end] != '|') {
if (row_text[end] == '\\' and end + 1 < row_text.len) {
end += 2;
} else {
end += 1;
}
}
🤖 Prompt for AI Agents
In `@src/md/render_blocks.zig` around lines 74 - 83, The comment above the inner
scan loop in render_row (the while loop using start, end, row_text, and
cell_index) incorrectly states it skips "escaped chars and code spans" even
though the loop only skips escaped characters; update the comment to reflect
current behavior (e.g., "Find cell end, skipping escaped chars") and remove any
mention of backtick/code-span handling so it matches the implemented logic in
that loop.

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 1

Caution

Some comments are outside the diff and can’t be posted inline due to platform limitations.

⚠️ Outside diff range comments (1)
test/integration/bun-types/bun-types.test.ts (1)

37-61: Replace Bun.$ shell pipelines with Bun.spawn using bunExe and bunEnv from harness.

The file uses Bun.$ for shell commands that spawn Bun processes (lines 41–45, 51–61, 109), which bypasses the harness guarantees and violates the test guidance to avoid shell commands. This prevents using the same Bun build being tested and disables debug logging control.

Import bunExe and bunEnv from harness and refactor:

  • Line 37: Replace await $mkdir -p ${BASE_FIXTURE_DIR}.quiet(); with await mkdir(BASE_FIXTURE_DIR, { recursive: true }); (already imported).
  • Lines 41–45: Use Bun.spawn with explicit args, cwd: BUN_TYPES_PACKAGE_ROOT, bunExe, and bunEnv; await exit code and handle stderr on failure.
  • Lines 51–61: Same as above for the multi-command pipeline (bun run build, bun pm pack, rm, mv, etc.); consider splitting into separate spawns or using a shell wrapper only if unavoidable.
  • Line 109: Use Bun.spawn with cwd: fixtureDir, explicit args, bunExe, and bunEnv.
🤖 Fix all issues with AI agents
In `@test/integration/bun-types/bun-types.test.ts`:
- Around line 312-315: The constants expectedEmptyInterfacesWhenNoDOM and
expectedEmptyInterfacesThatReactDeclareWhenNoDOM are defined after the
describe(...) block but referenced inside multiple typeTest(...) calls executed
during module evaluation, causing a TDZ; move (hoist) the constant declarations
so they appear before the describe(...) invocation (i.e., place the definitions
above the describe block) so typeTest(...) can access them synchronously, and
ensure any related helper/exported symbols are updated accordingly.

Comment on lines +312 to 315
typeTest("checks without lib.dom.d.ts", {
emptyInterfaces: expectedEmptyInterfacesWhenNoDOM,
diagnostics: [],
});

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

⚠️ Potential issue | 🔴 Critical

🧩 Analysis chain

🏁 Script executed:

cd /tmp && find . -name "bun-types.test.ts" 2>/dev/null | head -5

Repository: oven-sh/bun

Length of output: 37


🏁 Script executed:

git ls-files | grep -E "bun-types\.test\.ts"

Repository: oven-sh/bun

Length of output: 99


🏁 Script executed:

head -100 test/integration/bun-types/bun-types.test.ts

Repository: oven-sh/bun

Length of output: 3407


🏁 Script executed:

sed -n '300,330p' test/integration/bun-types/bun-types.test.ts

Repository: oven-sh/bun

Length of output: 1290


🏁 Script executed:

sed -n '780,840p' test/integration/bun-types/bun-types.test.ts

Repository: oven-sh/bun

Length of output: 1458


🏁 Script executed:

wc -l test/integration/bun-types/bun-types.test.ts

Repository: oven-sh/bun

Length of output: 103


🏁 Script executed:

sed -n '1,50p' test/integration/bun-types/bun-types.test.ts | tail -30

Repository: oven-sh/bun

Length of output: 870


🏁 Script executed:

grep -n "const expectedEmptyInterfacesWhenNoDOM" test/integration/bun-types/bun-types.test.ts

Repository: oven-sh/bun

Length of output: 122


🏁 Script executed:

grep -n "describe(\"@types/bun integration test\"" test/integration/bun-types/bun-types.test.ts

Repository: oven-sh/bun

Length of output: 106


🏁 Script executed:

sed -n '310,320p' test/integration/bun-types/bun-types.test.ts

Repository: oven-sh/bun

Length of output: 390


🏁 Script executed:

sed -n '310,340p' test/integration/bun-types/bun-types.test.ts

Repository: oven-sh/bun

Length of output: 1219


🏁 Script executed:

sed -n '790,810p' test/integration/bun-types/bun-types.test.ts

Repository: oven-sh/bun

Length of output: 494


🏁 Script executed:

sed -n '785,795p' test/integration/bun-types/bun-types.test.ts

Repository: oven-sh/bun

Length of output: 290


🏁 Script executed:

grep -n "expectedEmptyInterfacesThatReactDeclareWhenNoDOM\|expectedEmptyInterfacesWhenNoDOM" test/integration/bun-types/bun-types.test.ts | head -20

Repository: oven-sh/bun

Length of output: 862


Hoist expectedEmptyInterfacesWhenNoDOM and expectedEmptyInterfacesThatReactDeclareWhenNoDOM above the describe block.

These constants are accessed in multiple typeTest() calls within the describe() callback (lines 313, 337, 343, 428, 450, 474, 499, 510, 520, 530), but are defined at lines 794–929, after the describe block closes. Since describe() executes its callback synchronously during module evaluation, this creates a Temporal Dead Zone error and prevents the module from loading.

🛠️ Proposed fix
+const expectedEmptyInterfacesWhenNoDOM = new Set(["ThisType"]);
+
+const expectedEmptyInterfacesThatReactDeclareWhenNoDOM = new Set([
+  ...expectedEmptyInterfacesWhenNoDOM,
+  "Document",
+  "DataTransfer",
+  "StyleMedia",
+  "Element",
+  "DocumentFragment",
+  "HTMLElement",
+  "HTMLAnchorElement",
+  "HTMLAreaElement",
+  "HTMLAudioElement",
+  "HTMLBaseElement",
+  "HTMLBodyElement",
+  "HTMLBRElement",
+  "HTMLButtonElement",
+  "HTMLCanvasElement",
+  "HTMLDataElement",
+  "HTMLDataListElement",
+  "HTMLDetailsElement",
+  "HTMLDialogElement",
+  "HTMLDivElement",
+  "HTMLDListElement",
+  "HTMLEmbedElement",
+  "HTMLFieldSetElement",
+  "HTMLFormElement",
+  "HTMLHeadingElement",
+  "HTMLHeadElement",
+  "HTMLHRElement",
+  "HTMLHtmlElement",
+  "HTMLIFrameElement",
+  "HTMLImageElement",
+  "HTMLInputElement",
+  "HTMLModElement",
+  "HTMLLabelElement",
+  "HTMLLegendElement",
+  "HTMLLIElement",
+  "HTMLLinkElement",
+  "HTMLMapElement",
+  "HTMLMetaElement",
+  "HTMLMeterElement",
+  "HTMLObjectElement",
+  "HTMLOListElement",
+  "HTMLOptGroupElement",
+  "HTMLOptionElement",
+  "HTMLOutputElement",
+  "HTMLParagraphElement",
+  "HTMLParamElement",
+  "HTMLPreElement",
+  "HTMLProgressElement",
+  "HTMLQuoteElement",
+  "HTMLSlotElement",
+  "HTMLScriptElement",
+  "HTMLSelectElement",
+  "HTMLSourceElement",
+  "HTMLSpanElement",
+  "HTMLStyleElement",
+  "HTMLTableElement",
+  "HTMLTableColElement",
+  "HTMLTableDataCellElement",
+  "HTMLTableHeaderCellElement",
+  "HTMLTableRowElement",
+  "HTMLTableSectionElement",
+  "HTMLTemplateElement",
+  "HTMLTextAreaElement",
+  "HTMLTimeElement",
+  "HTMLTitleElement",
+  "HTMLTrackElement",
+  "HTMLUListElement",
+  "HTMLVideoElement",
+  "HTMLWebViewElement",
+  "SVGElement",
+  "SVGSVGElement",
+  "SVGCircleElement",
+  "SVGClipPathElement",
+  "SVGDefsElement",
+  "SVGDescElement",
+  "SVGEllipseElement",
+  "SVGFEBlendElement",
+  "SVGFEColorMatrixElement",
+  "SVGFEComponentTransferElement",
+  "SVGFECompositeElement",
+  "SVGFEConvolveMatrixElement",
+  "SVGFEDiffuseLightingElement",
+  "SVGFEDisplacementMapElement",
+  "SVGFEDistantLightElement",
+  "SVGFEDropShadowElement",
+  "SVGFEFloodElement",
+  "SVGFEFuncAElement",
+  "SVGFEFuncBElement",
+  "SVGFEFuncGElement",
+  "SVGFEFuncRElement",
+  "SVGFEGaussianBlurElement",
+  "SVGFEImageElement",
+  "SVGFEMergeElement",
+  "SVGFEMergeNodeElement",
+  "SVGFEMorphologyElement",
+  "SVGFEOffsetElement",
+  "SVGFEPointLightElement",
+  "SVGFESpecularLightingElement",
+  "SVGFESpotLightElement",
+  "SVGFETileElement",
+  "SVGFETurbulenceElement",
+  "SVGFilterElement",
+  "SVGForeignObjectElement",
+  "SVGGElement",
+  "SVGImageElement",
+  "SVGLineElement",
+  "SVGLinearGradientElement",
+  "SVGMarkerElement",
+  "SVGMaskElement",
+  "SVGMetadataElement",
+  "SVGPathElement",
+  "SVGPatternElement",
+  "SVGPolygonElement",
+  "SVGPolylineElement",
+  "SVGRadialGradientElement",
+  "SVGRectElement",
+  "SVGSetElement",
+  "SVGStopElement",
+  "SVGSwitchElement",
+  "SVGSymbolElement",
+  "SVGTextElement",
+  "SVGTextPathElement",
+  "SVGTSpanElement",
+  "SVGUseElement",
+  "SVGViewElement",
+  "Text",
+  "TouchList",
+  "WebGLRenderingContext",
+  "WebGL2RenderingContext",
+  "TrustedHTML",
+  "MediaStream",
+  "MediaSource",
+]);
+
 describe("@types/bun integration test", () => {
   describe("basic type checks", () => {
     typeTest("checks without lib.dom.d.ts", {
       emptyInterfaces: expectedEmptyInterfacesWhenNoDOM,
       diagnostics: [],
     });
   });
   ...
 });
 
-const expectedEmptyInterfacesWhenNoDOM = new Set(["ThisType"]);
-
-const expectedEmptyInterfacesThatReactDeclareWhenNoDOM = new Set([
-  ...expectedEmptyInterfacesWhenNoDOM,
-  ...
-]);
🤖 Prompt for AI Agents
In `@test/integration/bun-types/bun-types.test.ts` around lines 312 - 315, The
constants expectedEmptyInterfacesWhenNoDOM and
expectedEmptyInterfacesThatReactDeclareWhenNoDOM are defined after the
describe(...) block but referenced inside multiple typeTest(...) calls executed
during module evaluation, causing a TDZ; move (hoist) the constant declarations
so they appear before the describe(...) invocation (i.e., place the definitions
above the describe block) so typeTest(...) can access them synchronously, and
ensure any related helper/exported symbols are updated accordingly.

@dylan-conway
dylan-conway merged commit 1bfe5c6 into main Jan 29, 2026
4 of 23 checks passed
@dylan-conway
dylan-conway deleted the jarred/mdx branch January 29, 2026 04:24
Jarred-Sumner pushed a commit that referenced this pull request Feb 3, 2026
### What does this PR do?

I was looking at the [recent
support](#26440) for markdown and did
some benchmarking against
[bindings](https://github.com/just-js/lo/blob/main/lib/md4c/api.js) i
created for my `lo` runtime to `md4c`. In some cases, Bun is quite a bit
slower, so i did a bit of digging and came up with this change. It uses
`indexOfAny` which should utilise `SIMD` where it's available to scan
ahead in the payload for characters that need escaping.

In
[benchmarks](https://gist.github.com/billywhizz/397f7929a8920c826c072139b695bb68#file-results-md)
I have done this results in anywhere from `3%` to `~15%` improvement in
throughput. The bigger the payload and the more space between entities
the bigger the gain afaict, which would make sense.

### How did you verify your code works?

It passes `test/js/bun/md/*.test.ts` running locally. Only tested on
macos. Can test on linux but I assume that will happen in CI anyway?

## main


![bun-main](https://github.com/user-attachments/assets/8b173b34-1f20-4e52-bb67-bb8b7e5658f3)

## patched


![bun-patch](https://github.com/user-attachments/assets/26bb600c-234c-4903-8f70-32f167481156)
xhjkl pushed a commit to xhjkl/bun that referenced this pull request May 14, 2026
## Summary

- Port md4c (CommonMark-compliant markdown parser) from C to Zig under
`src/md/`
- Three output modes:
  - `Bun.markdown.html(input, options?)` — render to HTML string
- `Bun.markdown.render(input, callbacks?)` — render with custom
callbacks for each element
- `Bun.markdown.react(input, options?)` — render to a React Fragment
element, directly usable as a component return value
- React element creation uses a cached JSC Structure with
`putDirectOffset` for fast allocation
- Component overrides in `react()`: pass tag names as options keys to
replace default HTML elements with custom components
- GFM extensions: tables, strikethrough, task lists, permissive
autolinks, disallowed raw HTML tag filter
- Wire up `.md` as a bundler loader (via explicit `{ type: "md" }`)

## JavaScript API

### `Bun.markdown.html(input, options?)`

Renders markdown to an HTML string:

```js
const html = Bun.markdown.html("# Hello **world**");
// "<h1>Hello <strong>world</strong></h1>\n"

Bun.markdown.html("## Hello", { headingIds: true });
// '<h2 id="hello">Hello</h2>\n'
```

### `Bun.markdown.render(input, callbacks?)`

Renders markdown with custom JavaScript callbacks for each element. Each
callback receives children as a string and optional metadata, and
returns a string:

```js
// Custom HTML with classes
const html = Bun.markdown.render("# Title\n\nHello **world**", {
  heading: (children, { level }) => `<h${level} class="title">${children}</h${level}>`,
  paragraph: (children) => `<p>${children}</p>`,
  strong: (children) => `<b>${children}</b>`,
});

// ANSI terminal output
const ansi = Bun.markdown.render("# Hello\n\n**bold**", {
  heading: (children) => `\x1b[1;4m${children}\x1b[0m\n`,
  paragraph: (children) => children + "\n",
  strong: (children) => `\x1b[1m${children}\x1b[22m`,
});

// Strip all formatting
const text = Bun.markdown.render("# Hello **world**", {
  heading: (children) => children,
  paragraph: (children) => children,
  strong: (children) => children,
});
// "Hello world"

// Return null to omit elements
const result = Bun.markdown.render("# Title\n\n![logo](img.png)\n\nHello", {
  image: () => null,
  heading: (children) => children,
  paragraph: (children) => children + "\n",
});
// "Title\nHello\n"
```

Parser options can be included alongside callbacks:

```js
Bun.markdown.render("Visit www.example.com", {
  link: (children, { href }) => `[${children}](${href})`,
  paragraph: (children) => children,
  permissiveAutolinks: true,
});
```

### `Bun.markdown.react(input, options?)`

Returns a React Fragment element — use it directly as a component return
value:

```tsx
// Use as a component
function Markdown({ text }: { text: string }) {
  return Bun.markdown.react(text);
}

// With custom components
function Heading({ children }: { children: React.ReactNode }) {
  return <h1 className="title">{children}</h1>;
}
const element = Bun.markdown.react("# Hello", { h1: Heading });

// Server-side rendering
import { renderToString } from "react-dom/server";
const html = renderToString(Bun.markdown.react("# Hello **world**"));
// "<h1>Hello <strong>world</strong></h1>"
```

#### React 18 and older

By default, `react()` uses `Symbol.for('react.transitional.element')` as
the `$$typeof` symbol, which is what React 19 expects. For React 18 and
older, pass `reactVersion: 18`:

```tsx
const el = Bun.markdown.react("# Hello", { reactVersion: 18 });
```

### Component Overrides

Tag names can be overridden in `react()`:

```tsx
Bun.markdown.react(input, {
  h1: MyHeading,      // block elements
  p: CustomParagraph,
  a: CustomLink,      // inline elements
  img: CustomImage,
  pre: CodeBlock,
  // ... h1-h6, p, blockquote, ul, ol, li, pre, hr, html,
  //     table, thead, tbody, tr, th, td,
  //     em, strong, a, img, code, del, math, u, br
});
```

Boolean values are ignored (not treated as overrides), so parser options
like `{ strikethrough: true }` don't conflict with component overrides.

### Options

```js
Bun.markdown.html(input, {
  tables: true,              // GFM tables (default: true)
  strikethrough: true,       // ~~deleted~~ (default: true)
  tasklists: true,           // - [x] items (default: true)
  headingIds: true,          // Generate id attributes on headings
  autolinkHeadings: true,    // Wrap heading content in <a> tags
  tagFilter: false,          // GFM disallowed HTML tags
  wikiLinks: false,          // [[wiki]] links
  latexMath: false,          // $inline$ and $$display$$
  underline: false,          // __underline__ (instead of <strong>)
  // ... and more
});
```

## Architecture

### Parser (`src/md/`)

The parser is split into focused modules using Zig's delegation pattern:

| Module | Purpose |
|--------|---------|
| `parser.zig` | Core `Parser` struct, state, and re-exported method
delegation |
| `blocks.zig` | Block-level parsing: document processing, line
analysis, block start/end |
| `containers.zig` | Container management: blockquotes, lists, list
items |
| `inlines.zig` | Inline parsing: emphasis, code spans, HTML tags,
entities |
| `links.zig` | Link/image resolution, reference links, autolink
rendering |
| `autolinks.zig` | Permissive autolink detection (www, url, email) |
| `line_analysis.zig` | Line classification: headings, fences, HTML
blocks, tables |
| `ref_defs.zig` | Reference definition parsing and lookup |
| `render_blocks.zig` | Block rendering dispatch (code, HTML, table
blocks) |
| `html_renderer.zig` | HTML renderer implementing `Renderer` VTable |
| `types.zig` | Shared types: `Renderer` VTable, `BlockType`,
`SpanType`, `TextType`, etc. |

### Renderer Abstraction

Parsing is decoupled from output via a `Renderer` VTable interface:

```zig
pub const Renderer = struct {
    ptr: *anyopaque,
    vtable: *const VTable,

    pub const VTable = struct {
        enterBlock: *const fn (...) void,
        leaveBlock: *const fn (...) void,
        enterSpan:  *const fn (...) void,
        leaveSpan:  *const fn (...) void,
        text:       *const fn (...) void,
    };
};
```

Four renderers are implemented:
- **`HtmlRenderer`** (`src/md/html_renderer.zig`) — produces HTML string
output
- **`JsCallbackRenderer`** (`src/bun.js/api/MarkdownObject.zig`) — calls
JS callbacks for each element, accumulates string output
- **`ParseRenderer`** (`src/bun.js/api/MarkdownObject.zig`) — builds
React element AST with `MarkedArgumentBuffer` for GC safety
- **`JSReactElement`** (`src/bun.js/bindings/JSReactElement.cpp`) — C++
fast path for React element creation using cached JSC Structure +
`putDirectOffset`

## Test plan

- [x] 792 spec tests pass (CommonMark, GFM tables, strikethrough,
tasklists, permissive autolinks, GFM tag filter, wiki links, coverage,
regressions)
- [x] 114 API tests pass (`html()`, `render()`, `react()`,
`renderToString` integration, component overrides)
- [x] 58 GFM compatibility tests pass

```
bun bd test test/js/bun/md/md-spec.test.ts       # 792 pass
bun bd test test/js/bun/md/md-render-api.test.ts  # 114 pass
bun bd test test/js/bun/md/gfm-compat.test.ts     # 58 pass
```

🤖 Generated with [Claude Code](https://claude.com/claude-code)

---------

Co-authored-by: Claude <noreply@anthropic.com>
Co-authored-by: autofix-ci[bot] <114827586+autofix-ci[bot]@users.noreply.github.com>
Co-authored-by: Dylan Conway <dylan.conway567@gmail.com>
Co-authored-by: SUZUKI Sosuke <sosuke@bun.com>
Co-authored-by: robobun <robobun@oven.sh>
Co-authored-by: Claude Bot <claude-bot@bun.sh>
Co-authored-by: Kirill Markelov <kerusha.chubko@gmail.com>
Co-authored-by: Ciro Spaciari <ciro.spaciari@gmail.com>
Co-authored-by: Alistair Smith <hi@alistair.sh>
xhjkl pushed a commit to xhjkl/bun that referenced this pull request May 14, 2026
### What does this PR do?

I was looking at the [recent
support](oven-sh#26440) for markdown and did
some benchmarking against
[bindings](https://github.com/just-js/lo/blob/main/lib/md4c/api.js) i
created for my `lo` runtime to `md4c`. In some cases, Bun is quite a bit
slower, so i did a bit of digging and came up with this change. It uses
`indexOfAny` which should utilise `SIMD` where it's available to scan
ahead in the payload for characters that need escaping.

In
[benchmarks](https://gist.github.com/billywhizz/397f7929a8920c826c072139b695bb68#file-results-md)
I have done this results in anywhere from `3%` to `~15%` improvement in
throughput. The bigger the payload and the more space between entities
the bigger the gain afaict, which would make sense.

### How did you verify your code works?

It passes `test/js/bun/md/*.test.ts` running locally. Only tested on
macos. Can test on linux but I assume that will happen in CI anyway?

## main


![bun-main](https://github.com/user-attachments/assets/8b173b34-1f20-4e52-bb67-bb8b7e5658f3)

## patched


![bun-patch](https://github.com/user-attachments/assets/26bb600c-234c-4903-8f70-32f167481156)
liooil pushed a commit to liooil/poly that referenced this pull request Aug 7, 2026
### What does this PR do?

I was looking at the [recent
support](oven-sh/bun#26440) for markdown and did
some benchmarking against
[bindings](https://github.com/just-js/lo/blob/main/lib/md4c/api.js) i
created for my `lo` runtime to `md4c`. In some cases, Bun is quite a bit
slower, so i did a bit of digging and came up with this change. It uses
`indexOfAny` which should utilise `SIMD` where it's available to scan
ahead in the payload for characters that need escaping.

In
[benchmarks](https://gist.github.com/billywhizz/397f7929a8920c826c072139b695bb68#file-results-md)
I have done this results in anywhere from `3%` to `~15%` improvement in
throughput. The bigger the payload and the more space between entities
the bigger the gain afaict, which would make sense.

### How did you verify your code works?

It passes `test/js/bun/md/*.test.ts` running locally. Only tested on
macos. Can test on linux but I assume that will happen in CI anyway?

## main


![bun-main](https://github.com/user-attachments/assets/8b173b34-1f20-4e52-bb67-bb8b7e5658f3)

## patched


![bun-patch](https://github.com/user-attachments/assets/26bb600c-234c-4903-8f70-32f167481156)
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

7 participants